LAP: 언어-행동 사전 훈련은 제로샷 크로스-엠바디먼트 전송을 가능하게 한다
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
로봇 공학 분야의 오랜 목표는 새로운 로봇 플랫폼에 별도의 적응 과정 없이 즉시 적용 가능한 범용 정책을 개발하는 것입니다. 그러나 대규모 다중 플랫폼 사전 훈련에도 불구하고, 기존의 시각-언어-행동 모델(VLA)은 여전히 훈련 플랫폼에 강하게 의존하며, 일반적으로 비용이 많이 드는 미세 조정이 필요합니다. 본 연구에서는 언어-행동 사전 훈련(LAP)이라는 간단한 방법을 제시합니다. LAP은 저수준 로봇 동작을 자연어로 직접 표현하여, 행동 감독을 사전 훈련된 시각-언어 모델의 입력-출력 분포와 일치시킵니다. LAP은 학습된 토크나이저, 비용이 많이 드는 주석, 그리고 플랫폼별 특수한 아키텍처 설계가 필요하지 않습니다. LAP을 기반으로, 우리는 LAP-3B를 제안합니다. LAP-3B은 지금까지 알려진 바로는, 어떠한 플랫폼별 미세 조정 없이도 이전에 볼 수 없었던 로봇 플랫폼으로 상당한 제로샷 전송을 달성한 최초의 VLA 모델입니다. 여러 새로운 로봇 및 조작 작업에서, LAP-3B은 평균 50% 이상의 제로샷 성공률을 달성하며, 기존의 가장 강력한 VLA 모델보다 약 2배의 성능 향상을 보입니다. 또한, LAP은 효율적인 적응과 우수한 확장성을 가능하게 하며, 행동 예측과 시각 질의 응답(VQA)을 공유된 언어-행동 형식으로 통합하여 공동 훈련을 통해 추가적인 성능 향상을 제공합니다.
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.