EXPO-FT: 시각-언어-행동 모델의 효율적인 강화학습 미세 조정
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
로봇 공학 분야에서 새로운 작업을 효율적이고 안정적으로 학습하는 능력은 매우 중요한 과제입니다. 시각-언어-행동(VLA) 모델은 다양한 조작 작업에 대한 강력한 일반화 능력을 보여주었지만, 사전 훈련된 정책들은 실제 환경에서의 활용을 위한 충분한 신뢰성을 갖추지 못합니다. 강화 학습(RL) 미세 조정은 이러한 격차를 해소할 수 있는 유망한 방법이지만, 기존의 접근 방식은 사전 훈련된 지식을 완전히 활용하지 않거나, 효율적인 샘플 사용량과 성공률을 달성하지 못하는 VLA 모델 미세 조정을 수행합니다. 본 논문에서는 안정적이고 샘플 효율적인 RL 미세 조성을 통해 사전 훈련된 VLA 정책의 성능을 향상시키는 시스템인 EXPO-FT를 제시합니다. 저희 시스템은 스트링 라이트 연결, 플러그 연결하여 불 켜기, 당구공 넣기, 꽃병에 꽃 꽂기 등 다양한 어려운 조작 작업을 수행하며, 각 작업은 높은 정밀도, 동적 동작, 그리고 다양한 초기 상태에 대한 강건성을 요구합니다. 저희 시스템은 평균 19.1분의 로봇 데이터를 사용하여 평가된 모든 작업에서 완벽한 성능(30/30 성공)을 달성했으며, 기존의 강화 학습 기반 방법 및 VLA 미세 조정 방식보다 우수한 결과를 보였습니다. 저희는 로봇 공학 분야에서 VLA 모델의 RL 미세 조정을 더 널리 활용할 수 있도록 오픈 소스 코드를 공개합니다.
The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse manipulation tasks, yet pretrained policies consistently fall short of the reliability required for real-world deployment. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches either train from scratch without fully leveraging pretrained priors, or fine-tune VLAs without achieving the sample efficiency and success rates that practical deployment demands. We present EXPO-FT, a system for stable, sample-efficient RL finetuning of pretrained VLA policies that closes this gap. Our system solves a suite of challenging manipulation tasks, including routing string lights and inserting the plug to light it up, striking a pool ball into a pocket, and inserting a flower into a wine bottle, each requiring combinations of high precision, dynamic actions, and robustness to varied initial states. Our system achieves perfect task performance (30/30 successes) across all evaluated tasks within an average of 19.1 minutes of online robot data, outperforming both prior RL-from-scratch and VLA finetuning approaches. We release an open-source codebase with the aim of facilitating broader adoption of RL finetuning of VLA models in robotics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.