본능을 믿으세요: 신뢰도를 기반으로 하는 테스트 단계 강화 학습을 통한 시각-언어-행동 모델
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
강화 학습(RL)은 시각-언어-행동 모델(VLA)이 정적인 모방 학습의 한계를 넘어 발전하는 데 필수적인 기술이 되었습니다. 그러나 기존의 강화 학습 방법들은 일반적으로 외부 환경 피드백을 필요로 하며, 정책 업데이트를 안내하기 위해 미리 정의된 성공 신호에 의존합니다. 본 연구에서는 VLA 모델들이 유용한 내부 평가 능력을 가지고 있음을 보여줍니다. 특히 이산 행동 기반 VLA에서, 생성 신뢰도가 높은 경로일수록 성공할 가능성이 훨씬 높습니다. 이러한 관찰을 바탕으로, 우리는 아키텍처에 구애받지 않는 테스트 단계 강화 학습 프레임워크인 T^2VLA (Test-time VLA)를 제안합니다. T^2VLA는 VLA 모델이 자체적으로 정책 개선을 달성할 수 있도록 합니다. T^2VLA는 외부 보상에 의존하는 대신, 고신뢰 전문가 시연과 경로 수준의 유사성을 이용하여 내재적 보상 신호를 활용합니다. 또한, 우리는 탐색을 위한 로컬 가짜 전문가와 안정적인 학습을 위한 글로벌 전문가 풀을 동적으로 균형 있게 조정하는 신뢰도 기반 이중 전문가 부트스트래핑 메커니즘을 제안합니다. LIBERO 및 RoboTwin 벤치마크에서의 광범위한 실험 결과, T^2VLA는 기존의 지도 학습 방법보다 일관되게 우수한 성능을 보이며, 실제 보상을 사용한 강화 학습 수준에 근접하는 결과를 보여줍니다. 또한, T^2VLA는 OpenVLA-OFT 및 pi 시리즈를 포함한 다양한 VLA 파라다임에 적용될 수 있습니다.
Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. However, existing RL methods typically require external environmental feedback, relying on predefined success signals to guide policy updates. In this work, we show that VLA models possess useful internal evaluative capabilities: in discrete-action VLAs, trajectories with higher generation confidence are significantly more likely to succeed. Based on this observation, we introduce T^2VLA (Test-time VLA), an architecture-agnostic test-time RL framework that enables VLA models to achieve self-bootstrapping policy improvement. Instead of relying on external rewards, T^2VLA leverages trajectory-level similarity to high-confidence expert demonstrations as an intrinsic reward signal. In addition, we propose a Confidence-Driven Dual Expert Bootstrapping mechanism, which dynamically balances a Local Pseudo-Expert for exploration and a Global Expert Pool for training stability. Extensive experiments on the LIBERO and RoboTwin benchmarks show that T^2VLA consistently outperforms supervised baselines and approaches oracle RL performance with ground-truth rewards, achieving effective improvement without external reward feedback. Furthermore, T^2VLA adapts to distinct VLA paradigms, including both OpenVLA-OFT and the pi series.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.