BORA: 오프라인 강화 학습과 온라인 잔차 적응을 결합하여 실제 환경에서 정교한 조작 능력을 갖춘 VLA 모델 개발
BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models
Vision-Language-Action (VLA) 모델은 시각-언어 이해를 실제 로봇 조작에 적용하는 유망한 방법론으로 등장했습니다. 그러나 고차원적인 핸드 컨트롤과 누적되는 실행 오류로 인해, VLA 정책을 통해 정교한 조작 능력을 구현하는 것은 여전히 어려운 과제이며, 따라서 시각적으로 기반된 액션 생성과 물리적으로 신뢰할 수 있는 정교한 수행 간의 격차를 해소하기 위해 실제 환경에서의 추가 학습이 필수적입니다. 그러나 고차원적인 정교한 탐색은 종종 실제 환경에서 시간적 불일치, 샘플 효율성 저하 및 하드웨어 위험을 초래합니다. 이러한 문제점을 해결하기 위해, 우리는 실제 환경의 정교한 VLA 모델을 위한 오프라인-온라인 강화 학습 추가 학습 프레임워크인 BORA를 제안합니다. 오프라인 단계에서 BORA는 VLM의 인지 토큰과 액션 조각을 입력으로 받는 평가 모델(critic)을 구축합니다. 이러한 설계는 액션에 조건화된 가치 지침을 가능하게 하여, 평가 모델이 시각적 맥락뿐만 아니라 정교한 핸드 동작도 평가할 수 있도록 합니다. 이후 온라인 단계에서는 BORA가 기본 VLA 모델을 고정하고, 실제 환경에서의 실행 오류를 완화하고 오프라인에서 학습된 의도를 실제 물리적 환경 내에서 더욱 정확하게 수정하기 위해 경량의 인간-루프(HiL) 기반 조각별 잔차 적응 메커니즘을 도입합니다. BORA는 오프라인 평가 모델을 활용하고 개입 기반 보상을 사용하여 실행 오류를 효과적으로 수정하며, 사전 학습된 정책을 안정적인 기반으로 유지하면서 실제 환경의 물리적 변화에 적응합니다. 다섯 가지 복잡한 실제 정교한 조작 작업에 대한 광범위한 실험 결과는 BORA가 기존의 단순 모방 학습 및 전통적인 분리된 강화 학습 방법보다 훨씬 우수한 성능을 보이며, 표준 설정에서 평균 성공률이 33% 증가하고, 새로운 객체에 대한 일반화 성능이 최대 43% 향상된다는 것을 보여줍니다.
Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulation. However, dexterous manipulation remains challenging for VLA policies due to high-dimensional hand control and compounding execution errors, which makes real-world RL post-training essential for bridging the gap between visually grounded action generation and physically reliable dexterous execution. However, high-dimensional dexterous exploration often triggers temporal inconsistency, sample inefficiency and hardware risks in the real world. To address these challenges, we propose BORA, an offline-to-online RL post-training framework designed for real-world dexterous VLA models. In the offline phase, BORA constructs a critic that takes both the VLM's cognition tokens and action chunks as inputs. This design enables action-conditioned value guidance, allowing the critic to evaluate dexterous hand motions beyond visual context alone. During the subsequent online phase, BORA freezes the VLA base and introduces a lightweight, Human-in-the-Loop (HiL) chunk-wise residual adaptation mechanism to mitigate real-world execution errors and further correct the offline-learned intents within the actual physical environment. By inheriting the offline critic and employing intervention-driven rewards, BORA effectively corrects execution discrepancies and adapts to real-world physical variances while preserving the pretrained policy as a stable prior. Extensive evaluations across five complex real-world dexterous tasks demonstrate that BORA significantly outperforms pure imitation learning and traditional decoupled RL baselines, achieving a 33% absolute increase in average success rate under standard settings and up to a 43% improvement in unseen object generalization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.