2606.31846v1 Jun 30, 2026 cs.RO

Z-1: 시각-언어-행동 모델을 위한 효율적인 강화 학습

Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

Lang Cao
Lang Cao
Citations: 93
h-index: 5
Renhong Chen
Renhong Chen
Citations: 26
h-index: 2
Yitong Li
Yitong Li
Citations: 14
h-index: 2
Luyi Li
Luyi Li
Citations: 0
h-index: 0
Peng Wang
Peng Wang
Citations: 2
h-index: 1
Mofan Peng
Mofan Peng
Citations: 7
h-index: 1

시각-언어-행동(VLA) 모델은 언어 지침, 시각적 관찰 및 연속 제어를 연결하여 로봇 조작에 유망한 프레임워크를 제공합니다. 그러나 대부분의 기존 정책은 고정된 데모에서 이루어지는 행동 복제 또는 지도 미세 조정(SFT)에 의해 제한되며, 이는 정책 자체의 실패로부터 개선될 기회를 제한합니다. 본 논문에서는 흐름 기반 VLA 모델을 위한 강화 학습(RL) 후속 훈련 프레임워크인 Z-1을 제시합니다. Z-1은 공개된 RoboCasa 데모 데이터를 사용하여 SFT를 수행하고, $24$개의 표준 RoboCasa 작업에 대해 그룹 상대 정책 최적화(GRPO) 전략을 적용합니다. 온라인 최적화의 효율성과 안정성을 향상시키기 위해, Z-1은 공유 프리픽스 롤아웃 구성, 트리 구조 기반 트래젝토리 분기, 완료 인지 보상 조정 및 VLM과 액션 전문가의 선택적인 공동 훈련을 결합합니다. $24$개의 RoboCasa 작업에서 Z-1은 평균 성공률 $80.6%$를 달성했으며, 이는 SFT 초기화 대비 $13.2%$ 포인트 향상된 수치이며, 발표된 최첨단 모델보다 우수한 성능을 보입니다. 이러한 결과는 체계적인 GRPO 후속 훈련이 추가적인 비공개 데모 없이 흐름 기반 VLA 정책을 크게 개선할 수 있음을 보여줍니다.

Original Abstract

Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed demonstrations, which provides limited opportunity to improve from the policy's own failures. In this paper, we present Z-1, a reinforcement learning (RL) post-training framework for flow-based VLA models. Built on top of $π_{0.5}$, Z-1 uses only publicly released RoboCasa demonstrations for SFT and then applies a task-wise Group Relative Policy Optimization (GRPO) strategy across $24$ standard RoboCasa tasks. To improve the efficiency and stability of online optimization, Z-1 combines shared-prefix rollout construction, tree-structured trajectory branching, completion-aware reward calibration, and selective joint training of VLM and Action Expert. Across all $24$ RoboCasa tasks, Z-1 achieves an average success rate of $80.6\%$, improving over its SFT initialization by $13.2\%$ points and outperforms the published sota models. These results show that systematic GRPO post-training can substantially improve flow-based VLA policies without additional private demonstrations.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!