프리빌레지드 자기 증류를 통한 대조 학습 강화 정책 최적화
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
최근의 사전 학습된 거대 언어 모델(LLM) 발전은 검증 가능한 보상을 사용하는 강화 학습(RLVR) 또는 온-정책 자기 증류(OPSD)에 점점 더 의존하고 있습니다. OPSD는 밀집형, 로짓 수준의 지도 신호를 제공하지만, 자기 교사의 특권 정보로 인해 본질적으로 노출 편향 문제를 가지고 있으며, 이는 다중 턴 에이전트 환경에서 추론 경로의 수렴을 초래하고 명확한 최적화 방향을 상실하게 만듭니다. 이러한 문제점을 해결하기 위해, 우리는 대조 학습 관점에서 에이전트 OPSD를 재구성하는 Contrastive Reinforced Policy Optimization (CRPO)을 제안합니다. CRPO는 예측 엔트로피를 활용하여 긍정적인 위치(탐색적 탐색)와 부정적인 위치(노출 편향)를 구별하고, 그룹 단위의 대조 학습을 통해 신뢰할 수 있는 세밀한 최적화 신호를 유지합니다. 13개의 어려운 추론 및 심층 검색 벤치마크에 대한 광범위한 평가 결과, CRPO는 기존 강화 학습 및 자기 증류 방법보다 일관되게 우수한 성능을 보이며, 장기 상호 작용에서 학습 안정성과 일반화 능력을 크게 향상시킵니다.
Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.