HumanCLAW: 시각-언어 모델이 신체를 통해 행동할 수 있는가?
HumanCLAW: Can Vision-Language Models Act Through a Body?
시각-언어 모델(VLM)이 실제 신체를 통해 행동할 수 있는지 평가하는 것은 어려운 과제입니다. 행동의 결과는 VLM의 결정과 모터 제어가 결합된 형태이며, 작업 실패 시 VLM이 잘못된 선택을 했는지, 아니면 모터 컨트롤러가 단순히 실행에 실패했는지 (예: 균형을 잃고 넘어짐) 판단하기 어렵습니다. 본 연구에서는 행동 결정 과정을 저수준 실행 과정과 분리하는 평가 프레임워크인 HumanCLAW를 소개합니다. 이 시스템에서 VLM은 각 단계마다 간단한 동작 명령을 내리고, 해당 명령은 실제 물리적 제약 (중력, 충돌 등) 하에 신체 전체의 연속적인 움직임으로 변환되어 실행됩니다. 따라서 신체는 물리 세계에서 자유롭게 행동할 수 있으며, 실행 과정에서의 오류, 균형 문제 및 모터 오류는 결과 분석에서 제외됩니다. 측정 가능한 것은 모델의 행동 지능, 즉 신체가 다음에 어떤 동작을 수행해야 하는지에 대한 실시간 선택 능력입니다. 이 프레임워크를 기반으로 HumanCLAW-Bench라는 벤치마크를 구축했습니다. 이는 41개의 실내 환경에서 1,218개의 장기적인 시야각(egocentric) 탐색-상호작용 에피소드로 구성되어 있습니다. 최첨단 VLM 9개 모델을 테스트한 결과, 어떤 모델도 벤치마크를 해결하지 못했으며, 가장 성능이 좋은 모델의 성공률은 16.8%에 불과했습니다. 목표 인식을 못한 것이 문제인 것은 아닙니다. 현재 VLM들이 부족한 점은 '신체 자각'입니다. 즉, 자신의 신체의 위치를 파악하지 못하고, 목표 지점에 도달했는지, 또는 장애물에 부딪혔는지 등을 제대로 인식하지 못합니다.
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.