2607.29613v1 Jul 31, 2026 cs.RO

WCM: 시각-언어-행동 강화 학습을 위한 세계 모델

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Jingjing Gong
Jingjing Gong
Citations: 174
h-index: 6
Xipeng Qiu
Xipeng Qiu
Citations: 333
h-index: 9
Siyin Wang
Siyin Wang
Fudan University
Citations: 343
h-index: 9
Xiaopeng Yu
Xiaopeng Yu
Citations: 142
h-index: 2
Senyu Fei
Senyu Fei
Citations: 129
h-index: 2
Xianzhong Zhao
Xianzhong Zhao
Citations: 20
h-index: 3

시각-언어-행동(VLA) 모델의 강화 학습(RL) 후속 훈련은 로봇 조작 분야에서 뛰어난 성능을 보여주었습니다. RL 방법 중, 크리틱 기반 접근 방식은 주로 단일 프레임 관찰 또는 단일 프레임 VLM 백본 잠재 변수를 사용하여 값을 추정하는데, 이는 로봇 제어의 부분적으로 관찰 가능한 특성과 근본적인 불일치를 야기합니다. 관찰 기록을 통합하는 간단한 방법은 고차원 시각 공간으로 인해 지수적 복잡성을 초래하며, 순수한 스칼라-반환 회귀 방식은 시간 경과에 따른 동적 관계 학습에 필요한 충분한 감독 신호를 제공하지 못하기 때문에 실패합니다. 우리는 이러한 문제의 근본 원인이 명시적인 세계 모델링 목표가 없는 경우, 크리틱의 표현이 정확한 값 추정을 위해 필요한 시간 구조를 포착할 수 없다는 '상태 근사' 문제임을 밝혀냈습니다. 이를 해결하기 위해, 가벼운 LeJEPA 아키텍처 기반의 '세계 크리틱 모델(WCM)'을 제안합니다. WCM은 미래 잠재 상태를 예측하고 값을 동시에 추정함으로써, 크리틱의 표현이 단순히 스칼라 반환을 회귀하는 것이 아니라 시간 동적 관계를 명시적으로 학습하도록 훈련됩니다. WCM은 온-정책 및 오프-정책 훈련 파이프라인에 원활하게 통합되며, Pi0, Pi0.5, OpenVLA-OFT와 같은 최첨단 VLA 백본과 호환됩니다. 네 가지 벤치마크에서 총 149개의 작업에 대한 광범위한 실험 결과, WCM은 다양한 환경(분포 내 및 분포 외)에서 일관되게 최고 수준의 성능을 달성하며, 특히 일반화 능력이 뛰어남을 보여줍니다. 또한, OpenVLA-OFT와 Pi0.5를 사용하여 오프-정책 RL 방식으로 7개의 실제 로봇 조작 작업에 대해 WCM을 검증하여 다양한 환경에서 안정적인 배포가 가능함을 확인했습니다.

Original Abstract

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns. WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. We further validate WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!