2607.04927v1 Jul 06, 2026 cs.RO

DSWAM: 세밀한 로봇 조작을 위한 이중 시스템 기반의 세계 행동 모델

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

Zhangyuan Wang
Zhangyuan Wang
Citations: 58
h-index: 4
Taiyi Su
Taiyi Su
Citations: 396
h-index: 4
Kai Xie
Kai Xie
Citations: 0
h-index: 0
Jian Zhu
Jian Zhu
Citations: 0
h-index: 0
Jianjun Zhang
Jianjun Zhang
Citations: 29
h-index: 2
Tianbin Liu
Tianbin Liu
Citations: 0
h-index: 0
Zitai Huang
Zitai Huang
Citations: 5
h-index: 1
Chong Ma
Chong Ma
Citations: 36
h-index: 3
Youzhang He
Youzhang He
Citations: 2
h-index: 1
Tianjian Wang
Tianjian Wang
Citations: 12
h-index: 2
Hanyang Wang
Hanyang Wang
Citations: 492
h-index: 7
Weihao Ding
Weihao Ding
Citations: 0
h-index: 0
Yi Xu
Yi Xu
Citations: 142
h-index: 5

세계 행동 모델(WAM)은 비디오 기반의 세계 모델링을 활용하여 로봇 액션 학습에 대한 밀집적인 감독을 제공함으로써, 기존의 시각-언어-액션(VLA) 정책에 대한 유망한 대안을 제시합니다. 기존의 WAM은 물리적으로 실현 가능한 실행 능력이 뛰어나지만, 일반적으로 VLM 기반 VLA에서 사용되는 명시적인 언어 수준의 계획 인터페이스가 부족하여 전체적인 명령을 세분화하는 데 어려움을 겪습니다. 이러한 세분화는 가정 내 작업과 같이 복잡하고 다단계 목표를 포함하는 경우에 중요하며, 이때 전체적인 사용자 명령은 실행 가능한 세밀한 하위 작업 시퀀스로 변환되어야 합니다. 또한, 기존 시스템이 데이터, 로봇 플랫폼 및 작업 프로토콜 측면에서 다양하기 때문에, VLA와 WAM의 실제 로봇 성능을 공정하게 비교할 수 있는 연구가 부족합니다. 이러한 문제점과 더불어 VLA와의 체계적인 비교를 위해, 본 논문에서는 세밀한 로봇 조작을 위한 이중 시스템 기반의 세계 행동 모델인 DSWAM을 소개합니다. DSWAM은 기본 제어 경로로 System 1 WAM 실행기를 사용하며, 작업 분해가 유용한 경우에만 선택적으로 System 2 시각-언어 하위 작업 계획기를 활성화합니다. 계획기는 단기 시각 정보와 전역 작업 프롬프트를 기반으로 실행 가능한 하위 작업을 예측하고, WAM 실행기는 각 명령어 또는 하위 작업에 대한 세계 인식적인 행동을 생성합니다. 실행기는 액션 예측과 비디오 공동 학습을 통해 훈련되지만, 추론 과정에서는 명시적인 미래 비디오 생성을 사용하지 않고 직접적으로 액션 청크를 예측합니다. 이 실행 경로를 실제 로봇에서 활용 가능하도록 만들기 위해, TensorRT 가속, 비동기 실행 및 실시간 청킹(RTC) 기술을 통합하여 정책 쿼리가 로봇 제어를 방해하지 않도록 했습니다. 또한, VLA 정책과의 공정한 실제 로봇 비교를 위해, 동일한 로봇 플랫폼, 사전 학습 데이터, 사후 학습 데이터 및 평가 기준을 사용하여 DeMaVLA 환경에서 DSWAM을 구축하고 평가했습니다.

Original Abstract

World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs for decomposing coarse instructions. Such decomposition becomes important when household tasks involve complex multi-step goals, where coarse user commands need to be converted into sequences of fine-grained executable subtasks. Meanwhile, the field still lacks a fair real-robot comparison between VLA and WAM execution capabilities, since existing systems often differ in data, robot embodiments, and task protocols. To address both the decomposition gap and the need for a controlled WAM-VLA comparison, we introduce DSWAM, a Dual-System World Action Foundation Model for fine-grained robot manipulation. DSWAM keeps a System 1 WAM executor as the default control path and optionally activates a System 2 vision-language subtask planner only when task decomposition is useful. The planner predicts executable subtasks from short-term visual history and a global task prompt, while the WAM executor performs world-aware action generation for each instruction or subtask. The executor is trained with action prediction and video co-training, but inference directly predicts action chunks without explicit future video generation. To make this execution path practical on real robots, we further integrate TensorRT acceleration, asynchronous execution, and real-time chunking (RTC) so that policy queries do not block robot control. To provide a fair real-robot comparison with VLA policies, we build and evaluate DSWAM under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!