HERO: 환경 관찰을 통한 과거 정보 활용 강화 학습 및 에이전트 기반 자기 증류
HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
강화 학습은 일반적으로 경로의 최종 결과에 의존하여 다단계 에이전트의 능력을 향상시키므로, 각 중간 단계에서의 기여도를 파악하기 어렵습니다. 최근 연구에서는 온-정책 자기 증류 방법이, 특권 정보를 사용하여 밀집된 토큰 수준의 감독 신호를 제공하는 '자기 교사'를 통해 유망한 대안을 제시합니다. 본 연구는 다단계 환경에서 이 방식을 단순히 확장했을 때 관찰되는 성능 저하에 주목하며, 이는 성공적인 경로 또는 최종 결과와 같은 특권 정보와 학생 에이전트의 현재 의사 결정 상황 간의 불일치 때문이라고 분석합니다. 우리는 HERO라는 과거 정보를 활용하여 강화된 자기 증류 프레임워크를 제안합니다. HERO는 각 시행 후 완료된 상호 작용을 분석하여, 각 관찰 데이터를 원본 행동에 대한 실행 가능한 피드백(예: 필요성, 유효성 또는 실패 원인)을 담은 간결한 단계별 진단으로 변환합니다. TauBench 및 WebShop 환경에서 HERO는 기존의 환경-피드백 기반 자기 증류 방법 및 GRPO보다 작업 성공률을 향상시키고 불필요한 단계를 줄입니다. 특히, 학습 가능한 단계 수가 제한적이고 성공적인 시행이 드물어 GRPO가 약한 보상 대비 신호를 제공하는 경우에 HERO는 더욱 효과적입니다.
Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-policy self-distillation methods offer a promising alternative by converting privileged feedback into dense token-level supervision through a self-teacher. Our study is motivated by the unexpected performance degradation observed when naively extending this paradigm to multi-turn settings, which we attribute to a lack of alignment between privileged feedback, such as successful trajectories or terminal outcomes, and the student's current decision context. We introduce HERO, a hindsight-enhanced self-distillation framework that uses next environment observations as locally aligned feedback. After each rollout, HERO reflects on the completed interaction to convert each observation into a compact turn-level diagnosis, that captures actionable feedback about the original action such as its necessity, validity or failure cause. On TauBench and WebShop, HERO improves task success and reduces unnecessary turns over environment-feedback-only self-distillation and GRPO. It is especially effective under limited training turn budgets, where successful rollouts are rare and GRPO provides weak reward-contrast signals.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.