2607.28026v1 Jul 30, 2026 cs.LG

프리빌레지드 자기 증류를 통한 대조 학습 강화 정책 최적화

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Jianing Wang
Jianing Wang
Citations: 92
h-index: 6
Linsen Guo
Linsen Guo
Citations: 175
h-index: 6
Xingchen Liu
Xingchen Liu
Citations: 76
h-index: 4
Xingjian Wu
Xingjian Wu
Citations: 969
h-index: 12
Xuhan Zhu
Xuhan Zhu
Citations: 23
h-index: 3
Xiaoyu Li
Xiaoyu Li
Citations: 113
h-index: 6
Xuezhi Cao
Xuezhi Cao
Citations: 47
h-index: 3
Xunliang Cai
Xunliang Cai
Citations: 231
h-index: 8
Junlin Liu
Junlin Liu
Citations: 393
h-index: 6

최근의 사전 학습된 거대 언어 모델(LLM) 발전은 검증 가능한 보상을 사용하는 강화 학습(RLVR) 또는 온-정책 자기 증류(OPSD)에 점점 더 의존하고 있습니다. OPSD는 밀집형, 로짓 수준의 지도 신호를 제공하지만, 자기 교사의 특권 정보로 인해 본질적으로 노출 편향 문제를 가지고 있으며, 이는 다중 턴 에이전트 환경에서 추론 경로의 수렴을 초래하고 명확한 최적화 방향을 상실하게 만듭니다. 이러한 문제점을 해결하기 위해, 우리는 대조 학습 관점에서 에이전트 OPSD를 재구성하는 Contrastive Reinforced Policy Optimization (CRPO)을 제안합니다. CRPO는 예측 엔트로피를 활용하여 긍정적인 위치(탐색적 탐색)와 부정적인 위치(노출 편향)를 구별하고, 그룹 단위의 대조 학습을 통해 신뢰할 수 있는 세밀한 최적화 신호를 유지합니다. 13개의 어려운 추론 및 심층 검색 벤치마크에 대한 광범위한 평가 결과, CRPO는 기존 강화 학습 및 자기 증류 방법보다 일관되게 우수한 성능을 보이며, 장기 상호 작용에서 학습 안정성과 일반화 능력을 크게 향상시킵니다.

Original Abstract

Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!