2607.27180v1 Jul 29, 2026 cs.CV

HumanCLAW: 시각-언어 모델이 신체를 통해 행동할 수 있는가?

HumanCLAW: Can Vision-Language Models Act Through a Body?

Jiawei Gu
Jiawei Gu
National University of Singapore
Citations: 1,835
h-index: 9
Linjie Li
Linjie Li
Citations: 221
h-index: 5
Manling Li
Manling Li
Citations: 77
h-index: 3
Ranjay Krishna
Ranjay Krishna
Citations: 3,312
h-index: 22
Shuai Liu
Shuai Liu
Citations: 522
h-index: 5
Chengcheng Tang
Chengcheng Tang
Citations: 13
h-index: 2
Kairui Hu
Kairui Hu
Citations: 541
h-index: 4
Siyao Li
Siyao Li
Citations: 0
h-index: 0
Zekun Li
Zekun Li
Citations: 33
h-index: 3
Po-Chen Wu
Po-Chen Wu
Citations: 315
h-index: 8
Ivan Shugurov
Ivan Shugurov
Citations: 8
h-index: 2
Lingni Ma
Lingni Ma
Citations: 177
h-index: 4
Michael Zollhöfer
Michael Zollhöfer
Citations: 44
h-index: 3
Sizhe An
Sizhe An
Citations: 41
h-index: 2
Abhay Mittal
Abhay Mittal
Citations: 23
h-index: 3
Amy Zhao
Amy Zhao
Citations: 58
h-index: 4
Ziwei Liu
Ziwei Liu
Citations: 28
h-index: 1
Chuan Guo
Chuan Guo
Citations: 95
h-index: 5

시각-언어 모델(VLM)이 실제 신체를 통해 행동할 수 있는지 평가하는 것은 어려운 과제입니다. 행동의 결과는 VLM의 결정과 모터 제어가 결합된 형태이며, 작업 실패 시 VLM이 잘못된 선택을 했는지, 아니면 모터 컨트롤러가 단순히 실행에 실패했는지 (예: 균형을 잃고 넘어짐) 판단하기 어렵습니다. 본 연구에서는 행동 결정 과정을 저수준 실행 과정과 분리하는 평가 프레임워크인 HumanCLAW를 소개합니다. 이 시스템에서 VLM은 각 단계마다 간단한 동작 명령을 내리고, 해당 명령은 실제 물리적 제약 (중력, 충돌 등) 하에 신체 전체의 연속적인 움직임으로 변환되어 실행됩니다. 따라서 신체는 물리 세계에서 자유롭게 행동할 수 있으며, 실행 과정에서의 오류, 균형 문제 및 모터 오류는 결과 분석에서 제외됩니다. 측정 가능한 것은 모델의 행동 지능, 즉 신체가 다음에 어떤 동작을 수행해야 하는지에 대한 실시간 선택 능력입니다. 이 프레임워크를 기반으로 HumanCLAW-Bench라는 벤치마크를 구축했습니다. 이는 41개의 실내 환경에서 1,218개의 장기적인 시야각(egocentric) 탐색-상호작용 에피소드로 구성되어 있습니다. 최첨단 VLM 9개 모델을 테스트한 결과, 어떤 모델도 벤치마크를 해결하지 못했으며, 가장 성능이 좋은 모델의 성공률은 16.8%에 불과했습니다. 목표 인식을 못한 것이 문제인 것은 아닙니다. 현재 VLM들이 부족한 점은 '신체 자각'입니다. 즉, 자신의 신체의 위치를 파악하지 못하고, 목표 지점에 도달했는지, 또는 장애물에 부딪혔는지 등을 제대로 인식하지 못합니다.

Original Abstract

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!