숨겨진 움직임에서의 세계 모델 사고: Chain of World
Chain of World: World Model Thinking in Latent Motion
비전-언어-행동(VLA) 모델은 몸체 지능을 향한 유망한 경로이지만, 종종 시각적 역학의 기초가 되는 예측적이고 시간적 인과 관계 구조를 간과합니다. 세계 모델 기반 VLA는 미래 프레임을 예측하여 이러한 문제를 해결하지만, 불필요한 배경을 재구성하는 데 많은 계산 자원을 소비합니다. 잠재적 행동 기반 VLA는 프레임 간의 전환을 간결하게 인코딩하지만, 시간적으로 연속적인 동적 모델링 및 세계 지식이 부족합니다. 이러한 한계를 극복하기 위해, 우리는 세계 모델 기반의 시간적 추론과 분리된 잠재적 움직임 표현을 통합하는 새로운 "Chain of World" 패러다임인 CoWVLA (Chain-of-World VLA)를 제안합니다. 먼저, 사전 학습된 비디오 VAE가 잠재적 움직임 추출기로 작동하여 비디오 세그먼트를 구조 및 움직임 잠재 변수로 명시적으로 분리합니다. 그런 다음, 사전 학습 단계에서 VLA는 명령어와 초기 프레임을 통해 연속적인 잠재적 움직임 체인을 추론하고 세그먼트의 최종 프레임을 예측합니다. 마지막으로, 미세 조정 단계에서 이 잠재적 역학은 통합된 자기 회귀 디코더를 사용하여 희소한 주요 프레임과 행동 시퀀스를 함께 모델링함으로써 이산적인 행동 예측과 정렬됩니다. 이러한 설계는 시간적 추론 및 세계 지식의 세계 모델 이점을 유지하면서 잠재적 행동의 간결성과 해석성을 유지하여 효율적인 시각-운동 학습을 가능하게 합니다. 로봇 시뮬레이션 벤치마크에 대한 광범위한 실험 결과, CoWVLA는 기존의 세계 모델 및 잠재적 행동 기반 접근 방식보다 우수한 성능을 보이며 적당한 계산 효율성을 달성하여, 보다 효과적인 VLA 사전 학습 패러다임으로서의 잠재력을 보여줍니다. 프로젝트 웹사이트는 https://fx-hit.github.io/cowvla-io 에서 확인할 수 있습니다.
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.