WAM-TTT: 테스트 시간에 인간의 플레이를 관찰하여 월드 액션 모델을 제어하는 방법
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
로봇 기반 모델(RFM)을 새로운 작업 변형이나 사용자 선호 행동으로 유도하는 것은 여전히 어려운 과제이며, 종종 추가적인 로봇 데모, 작업별 미세 조정 또는 긴 문맥 조건 설정이 필요합니다. 본 논문에서는 WAM-TTT라는 테스트 시간 훈련 프레임워크를 제안합니다. 이 프레임워크는 원시 인간 비디오로부터 월드 액션 모델을 제어하는 방식으로 작동합니다. WAM-TTT는 인간 비디오를 모방할 수 있는 경로로 취급하는 대신, self-supervised 비디오 예측을 통해 동결된 WAM 내부의 경량 어댑티브 메모리에 흡수합니다. 이 메모리를 제어를 위해 활용하기 위해, 우리는 paired 인간-로봇 데이터를 사용하고 키-값 메모리 재구성 목적 함수를 통해 인간 데모와 로봇 행동을 정렬하는 메타 훈련 단계를 도입합니다. 테스트 시간에는 레이블이 없는 인간 비디오만 사용하여 메모리를 조정하며, 사전 훈련된 WAM은 그대로 유지됩니다. 이를 통해 로봇 액션, 인간 측면 주석 또는 작업별 미세 조정을 거치지 않고 효율적이고 재사용 가능한 제어가 가능하며, 동시에 기반 모델의 일반화 능력을 보존할 수 있습니다. 광범위한 실험 결과는 WAM-TTT가 다양한 조작 작업 및 일반화 환경에서 기존의 in-context 인간 비디오 조건부 방법보다 우수한 성능을 지속적으로 보여준다는 것을 입증합니다.
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.