평탄 정책을 넘어: 로봇 조작을 위한 임베디드 에이전트의 계층적 사후 학습
Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation
비전-언어-행동(VLA) 모델은 사전 학습된 비전-언어 모델을 활용하여 로봇 조작 분야에서 놀라운 성능을 보여주었습니다. 그러나 기존의 사후 학습 방법들은 주로 VLA 모델을 평탄한 정책으로 최적화하기 때문에, 작업 진행 과정을 명시적으로 모델링하고 안정적인 장기 계획 기반 조작을 수행하는 데 어려움이 있습니다. 계층적 접근 방식은 작업 분해를 도입하지만, 대부분 오프라인 시연 데이터를 활용한 지도 학습에 의존하며 온라인 상호 작용을 통해 실행 성능을 효과적으로 향상시키지 못합니다. 이러한 한계를 해결하기 위해 우리는 고수준 작업 계획과 저수준 행동 실행을 분리하는 계층적 사후 학습 프레임워크인 Hierarchical Robotic Control (HiRoC)을 제안합니다. 플래너는 복잡한 작업을 실행 가능한 하위 목표로 분해하여 명시적인 의미론적 지침을 제공하고, 실행기는 강화 학습을 통해 하위 목표에 조건화된 행동 생성 능력을 지속적으로 향상시킵니다. 두 모듈 간의 효과적인 협업을 가능하게 하기 위해, 우리는 강화 학습 전에 실행기가 플래너가 생성한 하위 목표와 일치하도록 조정하여 계획과 실행 사이의 분포 불일치를 완화합니다. 다양한 로봇 조작 벤치마크에서 수행된 광범위한 실험 결과는 HiRoC가 강력한 기준 모델보다 꾸준히 우수한 성능을 보임을 입증했습니다. 종합적인 분석은 계층적 사후 학습의 효과와 각 핵심 구성 요소의 기여를 더욱 검증합니다.
Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-training methods predominantly optimize VLA models as flat policies, making it difficult to explicitly model task progression and perform robust long-horizon manipulation. Although hierarchical approaches introduce task decomposition, they mainly rely on supervised learning from offline demonstrations and cannot effectively improve execution through online interaction. To address this limitation, we propose Hierarchical Robotic Control (HiRoC), a hierarchical post-training framework that decouples high-level task planning from low-level action execution. The planner decomposes complex tasks into executable subgoals to provide explicit semantic guidance, while the executor continuously improves subgoal-conditioned action generation through reinforcement learning. To enable effective collaboration between the two modules, we further align the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution. Extensive experiments across diverse robotic manipulation benchmarks demonstrate that HiRoC consistently outperforms strong baselines. Comprehensive analyses further validate the effectiveness of hierarchical post-training and the contribution of each key component.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.