2608.09730v1 Aug 10, 2026 cs.CV

월드 토큰(World Tokens): 학습 시간 세계 모델링을 통한 임베디드 정책 강화

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Bo Yuan
Bo Yuan
Citations: 25
h-index: 2
Longteng Guo
Longteng Guo
Citations: 1,611
h-index: 18
Qu Tang
Qu Tang
Citations: 0
h-index: 0
Benhui Zhuang
Benhui Zhuang
Citations: 38
h-index: 3
Xue Yu
Xue Yu
Citations: 0
h-index: 0
Junlan Feng
Junlan Feng
Citations: 2
h-index: 1

비전-언어-행동(VLA) 모델은 임베디드 정책에 널리 사용되는 패러다임입니다. 이들은 효율적인 폐루프 제어에 뛰어나지만, 작업이 진행됨에 따라 물리적 장면이 어떻게 변화하는지를 명시적으로 모델링하지 않습니다. 최근 등장한 세계-행동 모델(WAM)은 사전 학습된 비디오 세계 모델을 활용하여 시공간적 변화를 포착하지만, 미래 예측 또는 대규모 비디오 백본을 제어 루프에 유지하면 추론 비용이 크게 증가합니다. 본 논문에서는 World Adapter를 중심으로 구축된 임베디드 정책 아키텍처인 World Tokens을 소개합니다. World Adapter는 시각-언어 이해, 세계 역학 모델링 및 행동 생성을 연결합니다. 이는 학습 중에 세계 모델링을 사용하여 행동 정책을 향상시키는 동시에 효율적인 배포를 유지합니다. 구체적으로, World Adapter는 VLM(Vision-Language Model) 특징을 고정된 크기의 월드 토큰으로 변환하며, 이 토큰은 공동으로 미세 조정된 미래 비디오 디노이저에 대한 조건을 제공하고 동시에 행동 전문가의 유일한 시각-언어 컨텍스트 역할을 합니다. 이러한 공유 조건 덕분에 미래 비디오 디노이징에서 발생하는 기울기가 직접적으로 행동 예측에 사용되는 표현을 형성할 수 있으며, 전용 라우팅은 정책이 해당 표현을 우회하는 것을 방지합니다. 배포 시에는 세계 모델 분기를 제거하여 VLM, World Adapter 및 행동 전문가만 남게 되며, 온라인 비디오 모델 추론은 수행되지 않습니다. 20억 개의 파라미터 백본과 임베디드 액션 사전 학습 없이도 World Tokens은 LIBERO에서 높은 성능을 보이며, SIMPLER에서 보고된 최고 평균 성능을 달성하고, 동일한 행동 전용 기준선보다 실제 환경에서의 R1 Pro 성공률을 크게 향상시키며, 각 행동 단위를 VLA 수준의 지연 시간으로 생성합니다.

Original Abstract

Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert's sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!