CVPO: 가치-분산 적응 및 동적 커리큘럼 학습을 통한 LLM 강화 학습 추론 성능 향상
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
강화 학습(RL)은 대규모 언어 모델(LLM)의 추론 능력을 향상시키는 효과적인 방법으로 부상했습니다. 그러나 기존 방법들은 생성된 답변 경로에 대한 피드백의 정확성이 부족하고, 문제 난이도 변화 현상이 나타나는 단점이 있습니다. 이러한 문제를 해결하기 위해, 본 논문에서는 커리큘럼 기반 가치-분산 정책 최적화(Curriculum-guided Value-Variance Policy Optimization, CVPO)를 제안합니다. 응답 경로 수준에서, 토큰 단위의 가치-분산이 탐색 강도와 상관관계가 있음을 발견했습니다. 이론적 분석 결과, 이 분산은 정책 업데이트 크기를 제한하는 것으로 나타났습니다. 우리는 추정된 경로의 가치-분산을 사용하여 생성 과정에서의 고유한 무작위성을 정량화합니다. 이를 바탕으로, 다양한 보상 유형에 대한 분산 인지형 장점 조정 메커니즘을 설계했습니다. 질문 수준에서는, 문제 난이도에 맞춰 학습 단계를 조정하는 동적 커리큘럼 가중치 방법을 도입하여 모델이 각 학습 단계에서 현재 능력과 일치하는 작업에 집중하도록 돕습니다. 실험 결과, 제안된 방법은 VAPO와 같은 강력한 가치 기반 기법보다 우수한 성능을 보였습니다. 이는 다양한 수학 문제에서 언어 모델의 더욱 정확하고 안정적인 추론 능력을 가능하게 합니다.
Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. To address these challenges, we propose CVPO - Curriculum-guided Value-Variance Policy Optimization. At the response trajectory level, we find that token-level value-variance correlates with exploration intensity. Our theoretical analysis shows this variance bounds policy update magnitude. We then use the estimated trajectory value-variance to quantify the intrinsic randomness in generation. Based on this, we design a variance-aware advantage adjustment mechanism for different reward types. At the question level, we introduce a dynamic curriculum weighting method that adapts to question difficulty. This helps the model focus on tasks matched to its current ability during each training stage. Experimental results show our method outperforms strong value-based baselines like VAPO. It achieves better performance and stronger exploration, enabling more accurate and robust reasoning in language models across various math tasks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.