FORCE: 값 기반 초기 워밍업과 자기 증류를 통한 효율적인 VLA 강화 학습 미세 조정
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation
비전-언어-행동(VLA) 모델은 종종 최적이 아닌 데이터로 인해 발생하는 모방의 한계에 제약을 받습니다. 강화 학습(RL)을 통한 미세 조정은 이러한 제한을 극복할 수 있지만, 샘플 효율성이 매우 낮은 것으로 알려져 있습니다. 이 문제는 두 가지 핵심 문제에서 비롯됩니다: (1) 불안정한 Q 함수로 인한 초기 학습 오류, 그리고 (2) 저품질 탐색 데이터로 인한 비효율적인 정책 업데이트, 이는 종종 비용이 많이 드는 인간 개입을 필요로 합니다. 저희는 FORCE라는 3단계 프레임워크를 제안합니다. FORCE는 두 가지 문제를 해결하여 미세 조정을 안정화합니다. 먼저, 값 기반 초기 워밍업 단계를 도입하여 온-정책 시뮬레이션을 활용하고 Q 함수의 분포 변화를 완화합니다. 그 후, 온라인 단계에서 이 보정된 Q 함수는 정책 자체의 행동 제안과 전문가 데이터 모두를 필터링하는 역할을 합니다. 이를 통해 높은 가치를 가진 행동만 정책 업데이트에 사용됩니다. 저희는 다양한 시뮬레이션 및 실제 환경 작업에서 FORCE를 평가한 결과, 성공률이 79% 향상되었으며, 기존 RL 방법보다 10% 더 우수한 성능을 보였습니다. 또한, 학습 시간을 32.5% 단축했습니다. 특히, 일반적인 성공률 저하 문제를 완화하고 인간 개입 없이 강력한 성능을 달성하며, 능동적이고 자율적인 로봇 에이전트를 개발하는 데 중요한 진전을 이루었습니다.
Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit, it is notoriously sample inefficient. This challenge arises from two core issues: (1) catastrophic initial unlearning due to an unstable Q-function and (2) inefficient policy updates caused by low-quality exploration data, often forcing a reliance on costly human interventions. We introduce FORCE, a 3-stage framework that stabilizes fine-tuning by tackling both issues. FORCE first incorporates a Value-Calibrated Warm-Up phase, utilizing on-policy rollouts to mitigate the distributional shift of the Q-function. Subsequently, during the online stage, this calibrated Q-function acts as a filter for both the policy's own action proposals and expert data, ensuring only high-value actions are used for the policy update. We evaluate FORCE on various simulation and real-world tasks, and the result shows that FORCE achieves a 79% absolute improvement in success rates and outperform prior RL methods by 10%, while accelerating training by 32.5%. Critically, it mitigates the common success rate drop and achieves this robust performance without human intervention, marking a significant step towards deploying capable and autonomous robotic agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.