2607.02466v1 Jul 02, 2026 cs.RO

행동을 배우기 전에 움직임을 배우는 것: VLA 모델을 위한 작업 독립적인 사전 학습

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Jingjing Gong
Jingjing Gong
Citations: 174
h-index: 6
Xipeng Qiu
Xipeng Qiu
Citations: 333
h-index: 9
Siyin Wang
Siyin Wang
Fudan University
Citations: 343
h-index: 9
Junhao Shi
Junhao Shi
Fudan University
Citations: 160
h-index: 5
Xiaopeng Yu
Xiaopeng Yu
Citations: 142
h-index: 2
Li Ji
Li Ji
Citations: 137
h-index: 3

비전-언어-액션(VLA) 모델은 일반적으로 전문가의 시연 데이터 부족으로 인해 성능에 제한이 있습니다. 이러한 시연 데이터는 관찰, 지시 및 행동의 3가지 요소로 구성되며, 대규모 수집에는 상당한 비용이 소요됩니다. 본 연구에서는 이 문제점이 물리적 능력 습득(어떻게 움직이는가)과 의미적 정렬(무엇을 하는가)이라는 두 가지 별개의 학습 목표를 혼동하기 때문에 발생한다고 주장합니다. 특히, 후자는 언어적인 감독 정보만을 필요로 합니다. 이러한 분해 가설에 기반하여, 본 연구에서는 작업 독립적인 사전 학습(TAP)이라는 2단계 프레임워크를 제안합니다. TAP은 먼저 저렴하고 레이블이 없는 상호 작용 데이터(예: 버려진 비목표 경로 및 자율 로봇 플레이)로부터 전이 가능한 운동적 선행 지식을 자기 지도 Inverse Dynamics 방식을 통해 학습합니다. 이어서, 가벼운 2단계에서는 최소한의 전문가 데이터를 사용하여 이러한 선행 지식을 언어와 연결합니다. SIMPLER 벤치마크에서 TAP은 100만 건 이상의 전문가 시연 데이터로 학습된 모델과 유사한 성능을 보이면서도 훨씬 적은 양의 레이블된 데이터를 사용하며, 기존의 행동 복제 방식보다 10% 절대로 성능 향상을 달성합니다. 실제 환경인 WidowX 플랫폼에서 TAP은 카메라 변화에도 25%의 성공률을 유지하는 반면, 인터넷 규모의 기본 모델들은 0%로 떨어지는 것을 보여줍니다. 이는 작업 독립적인 사전 학습이 강력하고 전이 가능한 물리적 표현을 생성하며, Embodied AI 분야에 있어 확장 가능한 해결책을 제시한다는 것을 입증합니다.

Original Abstract

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale. We argue that this bottleneck stems from conflating two distinct learning objectives: acquiring physical competence (how to move) and acquiring semantic alignment (what to do). Crucially, only the latter requires language supervision. Building on this Decomposition Hypothesis, we propose Task-Agnostic Pretraining (TAP), a two-stage framework that first learns transferable motor priors from cheap, unlabeled interaction data -- including discarded off-task trajectories and autonomous robot play -- via a self-supervised Inverse Dynamics objective. A lightweight second stage then grounds these priors in language using minimal expert data. On the SIMPLER benchmark, TAP matches models trained on over 1M expert trajectories while using orders of magnitude less labeled data, yielding a 10% absolute gain over standard behavior cloning. On a real-world WidowX platform, TAP retains 25% success under camera perturbations where internet-scale baselines collapse to 0%, demonstrating that task-agnostic pretraining produces robust, transferable physical representations and offers a scalable path forward for Embodied AI.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!