2602.10556v2 Feb 11, 2026 cs.RO

LAP: 언어-행동 사전 훈련은 제로샷 크로스-엠바디먼트 전송을 가능하게 한다

LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer

Lihan Zha
Lihan Zha
Citations: 160
h-index: 7
Asher Hancock
Asher Hancock
Citations: 84
h-index: 5
Mingtong Zhang
Mingtong Zhang
Citations: 1,668
h-index: 11
Tenny Yin
Tenny Yin
Citations: 70
h-index: 5
Yixuan Huang
Yixuan Huang
Citations: 88
h-index: 6
Dhruv Shah
Dhruv Shah
Citations: 92
h-index: 4
Allen Z. Ren
Allen Z. Ren
Citations: 1,069
h-index: 12
Anirudha Majumdar
Anirudha Majumdar
Citations: 1,292
h-index: 10

로봇 공학 분야의 오랜 목표는 새로운 로봇 플랫폼에 별도의 적응 과정 없이 즉시 적용 가능한 범용 정책을 개발하는 것입니다. 그러나 대규모 다중 플랫폼 사전 훈련에도 불구하고, 기존의 시각-언어-행동 모델(VLA)은 여전히 훈련 플랫폼에 강하게 의존하며, 일반적으로 비용이 많이 드는 미세 조정이 필요합니다. 본 연구에서는 언어-행동 사전 훈련(LAP)이라는 간단한 방법을 제시합니다. LAP은 저수준 로봇 동작을 자연어로 직접 표현하여, 행동 감독을 사전 훈련된 시각-언어 모델의 입력-출력 분포와 일치시킵니다. LAP은 학습된 토크나이저, 비용이 많이 드는 주석, 그리고 플랫폼별 특수한 아키텍처 설계가 필요하지 않습니다. LAP을 기반으로, 우리는 LAP-3B를 제안합니다. LAP-3B은 지금까지 알려진 바로는, 어떠한 플랫폼별 미세 조정 없이도 이전에 볼 수 없었던 로봇 플랫폼으로 상당한 제로샷 전송을 달성한 최초의 VLA 모델입니다. 여러 새로운 로봇 및 조작 작업에서, LAP-3B은 평균 50% 이상의 제로샷 성공률을 달성하며, 기존의 가장 강력한 VLA 모델보다 약 2배의 성능 향상을 보입니다. 또한, LAP은 효율적인 적응과 우수한 확장성을 가능하게 하며, 행동 예측과 시각 질의 응답(VQA)을 공유된 언어-행동 형식으로 통합하여 공동 훈련을 통해 추가적인 성능 향상을 제공합니다.

Original Abstract

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodiment pre-training, existing Vision-Language-Action models (VLAs) remain tightly coupled to their training embodiments and typically require costly fine-tuning. We introduce Language-Action Pre-training (LAP), a simple recipe that represents low-level robot actions directly in natural language, aligning action supervision with the pre-trained vision-language model's input-output distribution. LAP requires no learned tokenizer, no costly annotation, and no embodiment-specific architectural design. Based on LAP, we present LAP-3B, which to the best of our knowledge is the first VLA to achieve substantial zero-shot transfer to previously unseen robot embodiments without any embodiment-specific fine-tuning. Across multiple novel robots and manipulation tasks, LAP-3B attains over 50% average zero-shot success, delivering roughly a 2x improvement over the strongest prior VLAs. We further show that LAP enables efficient adaptation and favorable scaling, while unifying action prediction and VQA in a shared language-action format that yields additional gains through co-training.

10 Citations
1 Influential
6 Altmetric
42.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!