2606.26006v1 Jun 24, 2026 cs.RO

FORCE: 값 기반 초기 워밍업과 자기 증류를 통한 효율적인 VLA 강화 학습 미세 조정

FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation

Pengwei Wang
Pengwei Wang
Citations: 981
h-index: 17
Shanghang Zhang
Shanghang Zhang
Citations: 921
h-index: 16
Y. Lou
Y. Lou
Citations: 9
h-index: 2
Zhongyuan Wang
Zhongyuan Wang
Citations: 950
h-index: 17
Shuyi Zhang
Shuyi Zhang
Citations: 127
h-index: 4
Hongyang Cheng
Hongyang Cheng
Citations: 51
h-index: 3
Yichen Guo
Yichen Guo
Citations: 8
h-index: 1
Chuyao Fu
Chuyao Fu
Citations: 3
h-index: 1
Yaoxu Lyu
Yaoxu Lyu
Citations: 331
h-index: 7
Xiaojie Zhang
Xiaojie Zhang
Citations: 56
h-index: 4
Haoran Li
Haoran Li
Citations: 0
h-index: 0

비전-언어-행동(VLA) 모델은 종종 최적이 아닌 데이터로 인해 발생하는 모방의 한계에 제약을 받습니다. 강화 학습(RL)을 통한 미세 조정은 이러한 제한을 극복할 수 있지만, 샘플 효율성이 매우 낮은 것으로 알려져 있습니다. 이 문제는 두 가지 핵심 문제에서 비롯됩니다: (1) 불안정한 Q 함수로 인한 초기 학습 오류, 그리고 (2) 저품질 탐색 데이터로 인한 비효율적인 정책 업데이트, 이는 종종 비용이 많이 드는 인간 개입을 필요로 합니다. 저희는 FORCE라는 3단계 프레임워크를 제안합니다. FORCE는 두 가지 문제를 해결하여 미세 조정을 안정화합니다. 먼저, 값 기반 초기 워밍업 단계를 도입하여 온-정책 시뮬레이션을 활용하고 Q 함수의 분포 변화를 완화합니다. 그 후, 온라인 단계에서 이 보정된 Q 함수는 정책 자체의 행동 제안과 전문가 데이터 모두를 필터링하는 역할을 합니다. 이를 통해 높은 가치를 가진 행동만 정책 업데이트에 사용됩니다. 저희는 다양한 시뮬레이션 및 실제 환경 작업에서 FORCE를 평가한 결과, 성공률이 79% 향상되었으며, 기존 RL 방법보다 10% 더 우수한 성능을 보였습니다. 또한, 학습 시간을 32.5% 단축했습니다. 특히, 일반적인 성공률 저하 문제를 완화하고 인간 개입 없이 강력한 성능을 달성하며, 능동적이고 자율적인 로봇 에이전트를 개발하는 데 중요한 진전을 이루었습니다.

Original Abstract

Vision-Language-Action (VLA) models are often constrained by the imitation ceiling imposed by sub-optimal data. While Reinforcement Learning (RL) fine-tuning can surpass this limit, it is notoriously sample inefficient. This challenge arises from two core issues: (1) catastrophic initial unlearning due to an unstable Q-function and (2) inefficient policy updates caused by low-quality exploration data, often forcing a reliance on costly human interventions. We introduce FORCE, a 3-stage framework that stabilizes fine-tuning by tackling both issues. FORCE first incorporates a Value-Calibrated Warm-Up phase, utilizing on-policy rollouts to mitigate the distributional shift of the Q-function. Subsequently, during the online stage, this calibrated Q-function acts as a filter for both the policy's own action proposals and expert data, ensuring only high-value actions are used for the policy update. We evaluate FORCE on various simulation and real-world tasks, and the result shows that FORCE achieves a 79% absolute improvement in success rates and outperform prior RL methods by 10%, while accelerating training by 32.5%. Critically, it mitigates the common success rate drop and achieves this robust performance without human intervention, marking a significant step towards deploying capable and autonomous robotic agents.

1 Citations
0 Influential
8.5 Altmetric
43.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!