강화 학습을 위한 단계 전환 기반 밀도 보상 모델링
Stage-Transition Dense Reward Modeling for Reinforcement Learning
장기적인 로봇 조작 작업을 위한 강화 학습은 종종 희소하고 지연된 보상 때문에 어려움을 겪습니다. 또한, 수동으로 설계된 밀도 형태의 신호는 비용이 많이 들고 환경 및 객체 구성 변경에 취약합니다. 본 연구에서는 Stage-Transition Dense Reward (STDR)라는 시각적 보상 학습 프레임워크를 제안합니다. STDR은 비정형 전문가 영상을 활용하여, 강화 학습 에이전트의 처음부터 학습을 위한 논리적으로 타당한 밀도 형태의 보상을 생성합니다. STDR은 시맨틱 이해를 통해 데모 영상에서 작업의 단계 구조를 추론하고, 온라인 훈련 중에 다음과 같은 두 가지 상호 보완적인 학습 신호를 제공합니다: (i) 목표 지향적인 보상을 제공하는 단계 전환 피드백, 그리고 (ii) 각 단계를 완료하기 위한 세밀한 지침을 제공하는 단계 내 진행 피드백. 또한, 이상 감지 메커니즘과 그립 제어 모듈을 통합하여 견고성을 향상시키고 보상 조작을 방지합니다. MetaWorld, ManiSkill 및 Franka Kitchen에서 14가지 조작 작업에 대한 실험 결과, STDR은 다양한 기준 모델보다 샘플 효율성과 성공률을 지속적으로 개선하며, 여러 가지 어려운 작업에서는 수동으로 설계된 밀도 형태의 보상과 동등하거나 더 나은 성능을 보여줍니다. 실제 로봇 평가 결과, STDR은 성공적인 실행 시 안정적이고 진행 상황에 맞는 보상을 부여하는 반면, 실패 시에는 적절히 낮은 보상을 생성하여, 시각적 노이즈에 대한 견고성과 다양한 환경에서 더 잘 조정된 보상 할당을 나타냅니다.
Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stage structure from demonstrations, and delivers two complementary learning signals during online training: (i) stage-transition feedback that provides goal-directed reward, and (ii) within-stage progress feedback that supplies fine-grained guidance toward completing each stage. Furthermore, an out-of-distribution (OOD) detection mechanism and a grasping regulation module are integrated to enhance robustness and prevent reward hacking. Experiments on 14 manipulation tasks across MetaWorld, ManiSkill, and Franka Kitchen show that STDR consistently improves sample efficiency and success rates over multiple baselines, and matches or surpasses handcrafted dense rewards on several challenging tasks. Real-robot evaluations further indicate that STDR assigns stable, progress-aligned rewards on successful executions while producing appropriately low rewards for failures, suggesting robustness to visual noise and better-calibrated reward assignment across settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.