교차 로봇 플랫폼 간 조작을 위한 액션 사전 학습
Learning Action Priors for Cross-embodiment Robot Manipulation
대부분의 비전-언어-액션(VLA) 모델은 비전-언어 모델(VLM)을 기반으로 하며, 액션 모듈을 추가하고 전체 정책을 동시에 최적화합니다. 이러한 설계는 VLM으로부터 강력한 시각 및 언어적 사전 정보를 상속받지만, 액션 모듈은 물리적인 움직임을 거의 처음부터 학습해야 합니다. 그 결과, 정책은 명시적인 운동 사전 정보가 부족하여 초기 최적화 단계에서 시간적 액션 역학 및 교차 모드 정렬을 동시에 학습해야 하며, 이는 특히 교차 로봇 플랫폼 환경에서 더욱 어려운 과제가 됩니다. 본 연구에서는 VLA 정렬 전에 액션 모듈에 운동 사전 정보를 미리 학습시키는 방법을 제안합니다. 구체적으로, 액션 모듈에 교차 로봇 플랫폼 환경에서의 시간적 운동 구조를 부여하는 두 단계의 훈련 프레임워크를 도입했습니다. 1단계에서는 가벼운 플로우 매칭 기반 인코더-디코더 액션 모듈이 시각 또는 언어 정보를 처리하지 않고, 비조건화된 액션 경로 데이터만 사용하여 효율적으로 시간적 운동 구조를 학습합니다. 2단계에서는 이 학습된 사전 정보가 디코더 재사용 및 초기 단계의 잠재 변환 증류를 통해 VLA 훈련에 전달되어, 시각-언어 특징을 액션 임베딩 공간과 정렬하면서도 전체 정책의 미세 조정을 허용합니다. 또한, 훈련된 인코더는 상태-액션 히스토리를 단일 시간적 컨텍스트 토큰으로 압축하는 경량 히스토리 압축기로 사용되어, 거의 무시할 수 있는 비용으로 히스토리 기반 모델링을 가능하게 합니다. 시뮬레이션 및 실제 환경에서 13가지 다양한 교차 로봇 플랫폼 작업에 대한 광범위한 실험 결과는 제안된 방법의 효과를 입증합니다. 액션 사전 정보 없이 VLA 훈련하는 것과 비교하여, 우리의 모델은 더 빠른 수렴 속도, 높은 성공률, 그리고 데이터가 부족한 실제 환경 작업에서 현저히 향상된 성능을 보입니다. 더욱이, 1단계에서 사용되는 액션 데이터를 늘리면 더욱 일반화 가능한 운동 사전 정보를 얻을 수 있으며, 이는 VLA의 downstream 성능을 직접적으로 향상시킵니다.
Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy jointly. This design inherits strong visual and linguistic priors from the VLM, but leaves the action module to learn physical motion almost from scratch. As a result, the policy lacks an explicit motion prior, forcing early optimization to simultaneously discover temporal action dynamics and cross-modal alignment, a challenge further amplified in cross-embodiment settings. In this work, we propose to pretrain the action module with motion priors before cross-modal VLA alignment. Specifically, we introduce a two-stage training framework that equips the action module with cross-embodiment temporal motion structure before VLA training begins. In Stage~1, a lightweight flow-matching-based encoder-decoder action module efficiently learns temporal motion structure solely from unconditioned action trajectories, without processing visual or language tokens. In Stage~2, this learned prior is transferred to VLA training through decoder reuse and early-stage latent distillation, aligning visual-language features with the action embedding space while still allowing end-to-end policy refinement. In addition, the trained encoder serves as a compact history compressor, summarizing state-action histories into a single temporal context token for history-aware modeling at negligible cost. Extensive experiments across 13 diverse cross-embodiment tasks on both simulated and real-world platforms validate the effectiveness of our approach. Compared with VLA training without action priors, our model achieves faster convergence, higher success rates, and substantially stronger performance on data-scarce real-world tasks. Moreover, scaling up the action data in Stage~1 yields a more generalizable action prior that directly improves downstream VLA performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.