DSWAM: 세밀한 로봇 조작을 위한 이중 시스템 기반의 세계 행동 모델
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
세계 행동 모델(WAM)은 비디오 기반의 세계 모델링을 활용하여 로봇 액션 학습에 대한 밀집적인 감독을 제공함으로써, 기존의 시각-언어-액션(VLA) 정책에 대한 유망한 대안을 제시합니다. 기존의 WAM은 물리적으로 실현 가능한 실행 능력이 뛰어나지만, 일반적으로 VLM 기반 VLA에서 사용되는 명시적인 언어 수준의 계획 인터페이스가 부족하여 전체적인 명령을 세분화하는 데 어려움을 겪습니다. 이러한 세분화는 가정 내 작업과 같이 복잡하고 다단계 목표를 포함하는 경우에 중요하며, 이때 전체적인 사용자 명령은 실행 가능한 세밀한 하위 작업 시퀀스로 변환되어야 합니다. 또한, 기존 시스템이 데이터, 로봇 플랫폼 및 작업 프로토콜 측면에서 다양하기 때문에, VLA와 WAM의 실제 로봇 성능을 공정하게 비교할 수 있는 연구가 부족합니다. 이러한 문제점과 더불어 VLA와의 체계적인 비교를 위해, 본 논문에서는 세밀한 로봇 조작을 위한 이중 시스템 기반의 세계 행동 모델인 DSWAM을 소개합니다. DSWAM은 기본 제어 경로로 System 1 WAM 실행기를 사용하며, 작업 분해가 유용한 경우에만 선택적으로 System 2 시각-언어 하위 작업 계획기를 활성화합니다. 계획기는 단기 시각 정보와 전역 작업 프롬프트를 기반으로 실행 가능한 하위 작업을 예측하고, WAM 실행기는 각 명령어 또는 하위 작업에 대한 세계 인식적인 행동을 생성합니다. 실행기는 액션 예측과 비디오 공동 학습을 통해 훈련되지만, 추론 과정에서는 명시적인 미래 비디오 생성을 사용하지 않고 직접적으로 액션 청크를 예측합니다. 이 실행 경로를 실제 로봇에서 활용 가능하도록 만들기 위해, TensorRT 가속, 비동기 실행 및 실시간 청킹(RTC) 기술을 통합하여 정책 쿼리가 로봇 제어를 방해하지 않도록 했습니다. 또한, VLA 정책과의 공정한 실제 로봇 비교를 위해, 동일한 로봇 플랫폼, 사전 학습 데이터, 사후 학습 데이터 및 평가 기준을 사용하여 DeMaVLA 환경에서 DSWAM을 구축하고 평가했습니다.
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs for decomposing coarse instructions. Such decomposition becomes important when household tasks involve complex multi-step goals, where coarse user commands need to be converted into sequences of fine-grained executable subtasks. Meanwhile, the field still lacks a fair real-robot comparison between VLA and WAM execution capabilities, since existing systems often differ in data, robot embodiments, and task protocols. To address both the decomposition gap and the need for a controlled WAM-VLA comparison, we introduce DSWAM, a Dual-System World Action Foundation Model for fine-grained robot manipulation. DSWAM keeps a System 1 WAM executor as the default control path and optionally activates a System 2 vision-language subtask planner only when task decomposition is useful. The planner predicts executable subtasks from short-term visual history and a global task prompt, while the WAM executor performs world-aware action generation for each instruction or subtask. The executor is trained with action prediction and video co-training, but inference directly predicts action chunks without explicit future video generation. To make this execution path practical on real robots, we further integrate TensorRT acceleration, asynchronous execution, and real-time chunking (RTC) so that policy queries do not block robot control. To provide a fair real-robot comparison with VLA policies, we build and evaluate DSWAM under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.