Qwen-VLA: 작업, 환경 및 로봇 구현 방식에 따른 시각-언어-행동 모델의 통합
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
현재까지 몸체 지능 연구는 조작 또는 내비게이션과 같은 개별 작업에 특화된 모델을 통해 이루어져 왔으며, 이는 파편화된 기능과 작업, 환경 및 로봇 구현 방식 간의 제한적인 일반화로 이어집니다. 본 연구에서는 다양한 몸체 의사 결정 문제를 단일 시각-언어-행동 모델 내에서 통합할 수 있는지 조사합니다. 우리는 Qwen-VLA를 제안합니다. 이는 DiT 기반 액션 디코더를 통해 인식, 이해 및 추론을 넘어 연속적인 행동 및 경로 생성을 가능하게 하는 Qwen의 시각-언어 모델링 스택을 확장한 통합 몸체 기반 모델입니다. Qwen-VLA는 로봇 조작 경로, 인간의 1인칭 데모, 합성 시뮬레이션 데이터, 시각-언어 내비게이션 데이터, 경로 중심 감독 학습 및 보조 시각-언어 데이터를 포함한 다양한 데이터 소스를 활용한 대규모 공동 사전 훈련 방식으로 학습되었습니다. 여러 로봇 플랫폼을 지원하기 위해, 우리는 로봇별 텍스트 설명을 사용하여 현재 몸체 정보 및 제어 방식을 명시하는 '몸체 인식 프롬프트 조건부 입력' 방법을 도입했습니다. 또한, 조작, 내비게이션 및 경로 예측 작업을 통합된 행동 및 경로 예측 프레임워크로 구성하여, 로봇 형태, 작업 유형 및 환경에 따른 시각적 연관성, 공간 추론 및 연속적인 행동 생성을 가능하게 합니다. 조작, 내비게이션 및 경로 중심 벤치마크 실험 결과, 장면 레이아웃, 배경, 조명, 객체 구성 및 로봇 구현 방식의 변화에도 불구하고 일관된 다중 작업 성능과 일반화 능력을 보여줍니다. Qwen-VLA-Instruct는 LIBERO에서 97.9%, Simpler-WidowX에서 73.7%, RoboTwin-Easy/Hard에서 각각 86.1% 및 87.2%, R2R에서 69.0% OSR, RxR에서 59.6% SR, 실제 환경 ALOHA 실험에서 평균 76.9%의 일반화 성공률, 그리고 DOMINO 동적 조작에서 26.6%의 제로샷 성공률을 달성했습니다.
Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and auxiliary vision-language data. To support multiple robot platforms, we introduce embodiment-aware prompt conditioning, where robot-specific textual descriptions specify the current embodiment and control convention. We further cast manipulation, navigation, and trajectory prediction into a unified action-and-trajectory prediction framework, enabling transferable visual grounding, spatial reasoning, and continuous action generation across robot morphologies, task families, and environments. Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot embodiment. Qwen-VLA-Instruct achieves 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% OSR on R2R, 59.6% SR on RxR, 76.9% average OOD success in real-world ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.