통합 시각운동 제어 목표: 물리적 행동을 넘어선 VLA 모델의 학습
Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions
VLA(Vision-Language Action) 모델은 시각 및 언어 정보를 바탕으로 로봇의 행동을 예측하도록 학습됩니다. 이는 자연스러운 선택이지만, VLM(Vision-Language Model)이 풍부하고 고수준의 장면 및 목표 표현을 인코딩하는 반면, 로봇의 행동은 낮은 수준의 신호이며 제한적인 작업 구조를 갖는다는 불일치가 발생합니다. 본 연구에서는 정책 학습 대상을 변경함으로써, 오히려 아키텍처 설계 방식 자체를 바꾸지 않고도 더 좋고 효율적으로 학습된 정책을 얻을 수 있는지 질문합니다. 우리는 UVT(Unified Visuomotor Target)라는 통합 잠재적 예측 목표를 제안합니다. 이는 동기 제어와 시각적 장면 변화 정보를 동시에 인코딩하며, 아키텍처 변경이나 추가 데이터 없이 적용 가능합니다. 시뮬레이션 벤치마크 및 실제 양손 조작 작업을 포함한 두 가지 대표적인 VLA 시스템에 UVT를 적용한 결과, 학습 효율성, 최종 작업 성능 및 정책의 견고성이 향상되었으며, 특히 제한된 학습 예산과 어려운 환경 조건에서 더욱 큰 개선 효과가 나타났습니다. 관련 동영상 및 추가적인 질적 결과는 프로젝트 웹페이지(https://unified-visuomotor-targets.github.io/)에서 확인하실 수 있습니다.
VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representations of scenes and goals, while robot actions are low-level signals with limited task structure. We ask whether changing what the policy is trained to predict, rather than how it is architecturally designed, can yield better and more efficiently trained policies. We propose UVT (Unified Visuomotor Target), a unified latent prediction target that jointly encodes motor control and visual scene transition information, requiring no architectural changes and no additional data. Applied to two representative VLA systems across simulation benchmarks and real bimanual manipulation tasks, UVT improves training efficiency, final task performance, and policy robustness, with particularly strong gains under limited training budgets and challenging environmental conditions. Rollout videos and additional qualitative results are available at our project webpage: https://unified-visuomotor-targets.github.io/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.