CRL-VLA: 지속적 시각-언어-행동 학습
CRL-VLA: Continual Vision-Language-Action Learning
평생 학습은 개방형 환경의 체화된 에이전트(embodied agents)에게 필수적이며, 여기서 강화 학습 미세 조정은 시각-언어-행동(VLA) 모델이 환경과의 상호작용을 통해 정교한 조작 능력을 습득하게 하는 중요한 패러다임으로 부상했습니다. 따라서 지속적 강화 학습(CRL)은 평생 로봇 시나리오에 VLA 모델을 배포하기 위한 유망한 경로이지만, 안정성(기존 기술 유지)과 가소성(새로운 기술 습득) 간의 균형을 맞추는 것은 기존 방법들에게 여전히 어려운 과제로 남아 있습니다. 우리는 엄격한 이론적 경계를 바탕으로 VLA 모델의 지속적 사후 학습(post-training)을 위한 프레임워크인 CRL-VLA를 소개합니다. 우리는 안정성-가소성 트레이드오프를 정책 발산(policy divergence)으로 스케일링된 목표 조건부 어드밴티지(advantage) 크기와 연결하는 통합 성능 경계를 도출합니다. CRL-VLA는 비대칭 조절을 통해 이 딜레마를 해결합니다. 즉, 이전 작업에 대해서는 어드밴티지 크기를 제한하고, 새로운 작업에 대해서는 제어된 성장을 가능하게 합니다. 이는 새로운 목표 조건부 가치 공식화(GCVF)를 적용한 간단하지만 효과적인 이중 비평가(Dual-Critic) 아키텍처를 통해 구현되며, 여기서 고정된(frozen) 비평가는 의미론적 일관성을 유지하고, 학습 가능한 추정기는 적응을 주도합니다. LIBERO 벤치마크에 대한 실험은 CRL-VLA가 이러한 상충하는 목표들을 효과적으로 조화시켜, 망각 방지와 순방향 적응 모두에서 기준 모델들을 능가함을 보여줍니다.
Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation through environmental interaction. Thus, Continual Reinforcement Learning (CRL) is a promising pathway for deploying VLA models in lifelong robotic scenarios, yet balancing stability (retaining old skills) and plasticity (learning new ones) remains a formidable challenge for existing methods. We introduce CRL-VLA, a framework for continual post-training of VLA models with rigorous theoretical bounds. We derive a unified performance bound linking the stability-plasticity trade-off to goal-conditioned advantage magnitude, scaled by policy divergence. CRL-VLA resolves this dilemma via asymmetric regulation: constraining advantage magnitudes on prior tasks while enabling controlled growth on new tasks. This is realized through a simple but effective dual-critic architecture with novel Goal-Conditioned Value Formulation (GCVF), where a frozen critic anchors semantic consistency and a trainable estimator drives adaptation. Experiments on the LIBERO benchmark demonstrate that CRL-VLA effectively harmonizes these conflicting objectives, outperforming baselines in both anti-forgetting and forward adaptation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.