2607.01586v1 Jul 02, 2026 cs.CV

VLAFlow: 공동 학습 및 미래 잠재적 정렬을 통한 비전-언어-행동 모델의 통합 학습 프레임워크

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

Yan Xie
Yan Xie
Citations: 89
h-index: 3
Kun Zhan
Kun Zhan
Citations: 1,574
h-index: 16
Fengfa Li
Fengfa Li
Citations: 5
h-index: 1
H. Ji
H. Ji
Citations: 10
h-index: 2
Lei Ren
Lei Ren
Citations: 50
h-index: 3
Guoyang Xia
Guoyang Xia
Citations: 4
h-index: 1
Fangxiang Feng
Fangxiang Feng
Citations: 3,763
h-index: 16

최근 비전-언어-행동(VLA) 모델은 로봇 조작 분야에서 상당한 발전을 이루었지만, 다양한 로봇-데이터 사전 훈련 방식의 효과를 비교하기 어렵습니다. 기존 모델들은 종종 아키텍처, 데이터, 행동 공간 및 평가 프로토콜이 다르기 때문입니다. 본 논문에서는 VLA 모델 학습 목표를 체계적으로 비교할 수 있는 통합 프레임워크인 VLAFlow(Vision-Language-Action Flow)를 제시합니다. DROID, OpenX-Embodiment, OpenX-Augmented 및 RoboCOIN에서 약 5,000시간 분량의 데이터를 포함하는 이기종 로봇 데이터셋 OXEMix를 사용하여 동일한 pi0 스타일 아키텍처, 공유된 VLM 백본, 행동 전문가 및 14차원 행동 공간 하에서 다음 네 가지 방식을 평가합니다. 액션만을 사용하는 모델링(MindPI), 언어 기반의 공동 학습(MindLPI), 미래 잠재적 정렬(MindWPI) 및 이들의 결합(MindLWPI). LIBERO, LIBERO-Plus 및 SimplerEnv에서의 실험 결과, 액션만 사용하는 사전 훈련은 이기종 데이터에 민감하다는 것을 보여줍니다. 반면, 언어 기반의 지도는 비전-언어 일반화 능력을 유지하는 데 도움이 되며, 미래 잠재적 정렬은 상태 전이 및 행동 결과 모델링을 개선합니다. 두 가지 신호를 결합한 MindLWPI는 다양한 벤치마크에서 가장 안정적인 전체 성능을 보였습니다. 이러한 결과는 메타-행동 공간 관점을 제시하며, 언어와 미래 잠재적 표현은 상호 보완적인 중간 제약을 제공하여 이기종 행동 감독을 더욱 원활하고 전이 가능하게 만든다는 것을 시사합니다.

Original Abstract

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in architecture, data, action space, and evaluation protocol. We present VLAFlow (Vision-Language-Action Flow), a unified flow-matching framework for controlled comparison of VLA training objectives. Using a heterogeneous robot corpus, OXEMix, containing approximately 5,000 hours of data from DROID, OpenX-Embodiment, OpenX-Augmented, and RoboCOIN, we evaluate four paradigms under the same pi0-style architecture, shared VLM backbone, action expert, and 14-dimensional action space: action-only modeling (MindPI), language-supervised co-training (MindLPI), future latent alignment (MindWPI), and their combination (MindLWPI). Experiments on LIBERO, LIBERO-Plus, and SimplerEnv show that action-only pre-training is sensitive to heterogeneous data. In contrast, language supervision helps preserve vision-language generalization, while future latent alignment improves state-transition and action-outcome modeling. By combining both signals, MindLWPI achieves the most stable overall transfer performance across benchmarks. These results suggest a meta-action space view: language and future latent representations provide complementary intermediate constraints that make heterogeneous action supervision smoother and more transferable.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!