2605.30280v1 May 28, 2026 cs.RO

Qwen-VLA: 작업, 환경 및 로봇 구현 방식에 따른 시각-언어-행동 모델의 통합

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zixing Lei
Zixing Lei
Citations: 723
h-index: 6
Jian Guan
Jian Guan
Citations: 171
h-index: 5
Tong Zhang
Tong Zhang
Citations: 41
h-index: 4
Mingsheng Li
Mingsheng Li
Citations: 1,263
h-index: 2
Junyang Lin
Junyang Lin
Citations: 7,191
h-index: 10
Zhixuan Liang
Zhixuan Liang
Citations: 797
h-index: 10
Ji-lu Ye
Ji-lu Ye
Citations: 67
h-index: 3
Shuai Bai
Shuai Bai
Citations: 4
h-index: 1
Yuchong Sun
Yuchong Sun
Citations: 867
h-index: 9
Sicheng Xie
Sicheng Xie
Citations: 114
h-index: 4
Dayiheng Liu
Dayiheng Liu
Citations: 24,474
h-index: 27
Xuhong Huang
Xuhong Huang
Citations: 2
h-index: 1
Yitao Liu
Yitao Liu
Citations: 208
h-index: 3
Junhao Chen
Junhao Chen
Citations: 120
h-index: 6
Yingming Zheng
Yingming Zheng
Citations: 1
h-index: 1
Qiuyue Wang
Qiuyue Wang
Citations: 11
h-index: 1
Xintong Hu
Xintong Hu
Citations: 0
h-index: 0
Pei Lin
Pei Lin
Citations: 24
h-index: 3
Jiazhao Zhang
Jiazhao Zhang
Citations: 1,493
h-index: 19
Haoqi Yuan
Haoqi Yuan
Citations: 551
h-index: 12
G. Zhou
G. Zhou
Citations: 860
h-index: 8
Hang Yin
Hang Yin
Citations: 243
h-index: 5
Yebin Wang
Yebin Wang
Citations: 4
h-index: 1
Wujian Peng
Wujian Peng
Citations: 104
h-index: 4
Delin Chen
Delin Chen
Citations: 225
h-index: 8
Jingyang Fan
Jingyang Fan
Citations: 24
h-index: 2
Xianwei Zhuang
Xianwei Zhuang
Citations: 309
h-index: 12
Xinyu Zhou
Xinyu Zhou
Citations: 13
h-index: 2
Haoyang Li
Haoyang Li
Citations: 848
h-index: 7
An-Jen Chen
An-Jen Chen
Citations: 52
h-index: 4
Xuejing Liu
Xuejing Liu
Citations: 10,498
h-index: 5
Rui Chen
Rui Chen
Citations: 213
h-index: 5
Chenxu Lu
Chenxu Lu
Citations: 22
h-index: 3
Tao Yu
Tao Yu
Citations: 2
h-index: 1
Xiong-hui Chen
Xiong-hui Chen
Citations: 799
h-index: 4
Jie Zhang
Jie Zhang
Citations: 12
h-index: 2
Jing Zhou
Jing Zhou
Citations: 9
h-index: 1
Zhao Li
Zhao Li
Citations: 7
h-index: 2
Zhibo Yang
Zhibo Yang
Citations: 8,292
h-index: 26

현재까지 몸체 지능 연구는 조작 또는 내비게이션과 같은 개별 작업에 특화된 모델을 통해 이루어져 왔으며, 이는 파편화된 기능과 작업, 환경 및 로봇 구현 방식 간의 제한적인 일반화로 이어집니다. 본 연구에서는 다양한 몸체 의사 결정 문제를 단일 시각-언어-행동 모델 내에서 통합할 수 있는지 조사합니다. 우리는 Qwen-VLA를 제안합니다. 이는 DiT 기반 액션 디코더를 통해 인식, 이해 및 추론을 넘어 연속적인 행동 및 경로 생성을 가능하게 하는 Qwen의 시각-언어 모델링 스택을 확장한 통합 몸체 기반 모델입니다. Qwen-VLA는 로봇 조작 경로, 인간의 1인칭 데모, 합성 시뮬레이션 데이터, 시각-언어 내비게이션 데이터, 경로 중심 감독 학습 및 보조 시각-언어 데이터를 포함한 다양한 데이터 소스를 활용한 대규모 공동 사전 훈련 방식으로 학습되었습니다. 여러 로봇 플랫폼을 지원하기 위해, 우리는 로봇별 텍스트 설명을 사용하여 현재 몸체 정보 및 제어 방식을 명시하는 '몸체 인식 프롬프트 조건부 입력' 방법을 도입했습니다. 또한, 조작, 내비게이션 및 경로 예측 작업을 통합된 행동 및 경로 예측 프레임워크로 구성하여, 로봇 형태, 작업 유형 및 환경에 따른 시각적 연관성, 공간 추론 및 연속적인 행동 생성을 가능하게 합니다. 조작, 내비게이션 및 경로 중심 벤치마크 실험 결과, 장면 레이아웃, 배경, 조명, 객체 구성 및 로봇 구현 방식의 변화에도 불구하고 일관된 다중 작업 성능과 일반화 능력을 보여줍니다. Qwen-VLA-Instruct는 LIBERO에서 97.9%, Simpler-WidowX에서 73.7%, RoboTwin-Easy/Hard에서 각각 86.1% 및 87.2%, R2R에서 69.0% OSR, RxR에서 59.6% SR, 실제 환경 ALOHA 실험에서 평균 76.9%의 일반화 성공률, 그리고 DOMINO 동적 조작에서 26.6%의 제로샷 성공률을 달성했습니다.

Original Abstract

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder. Qwen-VLA is trained with a large-scale joint pretraining recipe over diverse data sources, including robotics manipulation trajectories, human egocentric demonstrations, synthetic simulation data, vision-and-language navigation data, trajectory-centric supervision, and auxiliary vision-language data. To support multiple robot platforms, we introduce embodiment-aware prompt conditioning, where robot-specific textual descriptions specify the current embodiment and control convention. We further cast manipulation, navigation, and trajectory prediction into a unified action-and-trajectory prediction framework, enabling transferable visual grounding, spatial reasoning, and continuous action generation across robot morphologies, task families, and environments. Experiments on manipulation, navigation, and trajectory-centric benchmarks show consistent multi-task performance and out-of-distribution generalization under variations in scene layout, background, lighting, object configuration, and robot embodiment. Qwen-VLA-Instruct achieves 97.9% on LIBERO, 73.7% on Simpler-WidowX, 86.1%/87.2% on RoboTwin-Easy/Hard, 69.0% OSR on R2R, 59.6% SR on RxR, 76.9% average OOD success in real-world ALOHA experiments, and 26.6% zero-shot success on DOMINO dynamic manipulation.

22 Citations
1 Influential
13.5 Altmetric
91.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!