Z-1: 시각-언어-행동 모델을 위한 효율적인 강화 학습
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
시각-언어-행동(VLA) 모델은 언어 지침, 시각적 관찰 및 연속 제어를 연결하여 로봇 조작에 유망한 프레임워크를 제공합니다. 그러나 대부분의 기존 정책은 고정된 데모에서 이루어지는 행동 복제 또는 지도 미세 조정(SFT)에 의해 제한되며, 이는 정책 자체의 실패로부터 개선될 기회를 제한합니다. 본 논문에서는 흐름 기반 VLA 모델을 위한 강화 학습(RL) 후속 훈련 프레임워크인 Z-1을 제시합니다. Z-1은 공개된 RoboCasa 데모 데이터를 사용하여 SFT를 수행하고, $24$개의 표준 RoboCasa 작업에 대해 그룹 상대 정책 최적화(GRPO) 전략을 적용합니다. 온라인 최적화의 효율성과 안정성을 향상시키기 위해, Z-1은 공유 프리픽스 롤아웃 구성, 트리 구조 기반 트래젝토리 분기, 완료 인지 보상 조정 및 VLM과 액션 전문가의 선택적인 공동 훈련을 결합합니다. $24$개의 RoboCasa 작업에서 Z-1은 평균 성공률 $80.6%$를 달성했으며, 이는 SFT 초기화 대비 $13.2%$ 포인트 향상된 수치이며, 발표된 최첨단 모델보다 우수한 성능을 보입니다. 이러한 결과는 체계적인 GRPO 후속 훈련이 추가적인 비공개 데모 없이 흐름 기반 VLA 정책을 크게 개선할 수 있음을 보여줍니다.
Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and continuous control. However, most existing policies remain limited by behavior cloning or supervised fine-tuning (SFT) from fixed demonstrations, which provides limited opportunity to improve from the policy's own failures. In this paper, we present Z-1, a reinforcement learning (RL) post-training framework for flow-based VLA models. Built on top of $π_{0.5}$, Z-1 uses only publicly released RoboCasa demonstrations for SFT and then applies a task-wise Group Relative Policy Optimization (GRPO) strategy across $24$ standard RoboCasa tasks. To improve the efficiency and stability of online optimization, Z-1 combines shared-prefix rollout construction, tree-structured trajectory branching, completion-aware reward calibration, and selective joint training of VLM and Action Expert. Across all $24$ RoboCasa tasks, Z-1 achieves an average success rate of $80.6\%$, improving over its SFT initialization by $13.2\%$ points and outperforms the published sota models. These results show that systematic GRPO post-training can substantially improve flow-based VLA policies without additional private demonstrations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.