프레임당 하나의 토큰: VLA 정책을 위한 월드 모델에서 시각적 대역폭 재고
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
비전-언어-액션(VLA) 모델은 장기적인 계획 수립을 위해 점점 더 많은 보조 월드 모듈에 의존하고 있지만, 사전 학습된 VLA 모델에 이러한 모듈을 어떻게 구성할 것인지는 여전히 해결해야 할 설계 문제입니다. 기존의 월드 모델 기반 VLA는 일반적으로 각 프레임의 시각 정보를 월드 모듈에 높은 시각적 대역폭으로 전달하고, 이 과정을 액션 예측의 부산물로 취급합니다. 고정된 백본에 대한 제한된 적응 예산을 고려할 때, 이는 프레임별 표현과 잠재적 액션 결합 모두에 대한 충분한 검토를 어렵게 만듭니다. 본 연구에서는 Adaptive Attention Pooling을 통해 각 뷰를 프레임당 하나의 의미 있는 토큰으로 압축하는 OneWM-VLA를 제안합니다. OneWM-VLA는 결과적인 잠재적 스트림과 액션 경로를 별도의 디코더를 통해 연결하는 것이 아니라 단일 플로우 매칭 목표를 통해 생성합니다. 실험 결과, 제안하는 방식에서는 프레임별 시각적 대역폭을 단일 토큰으로 줄여도 장기적인 성능 저하 없이 작동한다는 것을 확인했습니다. 1471만 개의 LoRA 파라미터로 $π_0$ (2B) 백본을 사용하여 학습한 OneWM-VLA는 MetaWorld~MT50에서 평균 성공률을 47.9%에서 61.3%로 향상시켰으며, LIBERO-Long에서는 85.2% (vs. $π_0$의 95.6%), 실제 Piper 팔을 사용한 장기 변형 작업인 Fold Cloth에서는 20.0% (vs. $π_0$의 60.0%)로 향상시켰습니다.
Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameterized on top of a pretrained VLA remains an open design question. Existing world-model-augmented VLAs typically pass the per-frame visual stream into the world module at high visual bandwidth and treat its rollout as a side product of action prediction; under a constrained adaptation budget on a frozen backbone, this leaves both the per-frame representation and the latent action coupling under-examined. We introduce OneWM-VLA, which compresses each view into a single semantic token per frame through an Adaptive Attention Pooling, and produces the resulting latent stream and the action trajectory under a single flow-matching objective rather than connecting them through a separate decoder. Empirically, we find that per-frame visual bandwidth can be reduced to a single token without compromising long-horizon performance under our setup. Trained with 14.71M LoRA parameters on a $π_0$ (2B) backbone, OneWM-VLA improves the average success rate from 47.9% to 61.3% on MetaWorld~MT50, reaches 95.6% on LIBERO-Long (vs.85.2% for $π_0$), and reaches 60.0% on the long-horizon deformable task Fold Cloth on a real Piper arm (vs.20.0% for $π_0$).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.