EvoRubrics: 적대적 공진화를 통한 LLM 강화 학습을 위한 동적인 평가 기준
EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning
평가 기준 기반 보상은 검증 가능한 답변이 없는 개방형 작업에서 강화 학습을 위한 해석 가능하고 세밀한 최적화 신호를 제공합니다. 그러나 미리 구성된 평가 기준은 훈련 과정 전반에 걸쳐 고정되어 있으므로, 정책의 변화와 근본적인 불일치를 야기합니다. 즉, 모델이 개선됨에 따라 고정된 기준은 점차 구별력을 잃게 되어 보상 포화 및 잠재적인 해킹으로 이어질 수 있습니다. 최근 동적 평가 기준 방법은 이러한 문제를 부분적으로 해결하지만, 외부 최첨단 모델이나 정답을 사용하고, 평가 기준을 조잡한 수준에서만 업데이트합니다. 우리는 정책 LLM과 평가 기준 생성기가 각 훈련 단계 내에서 적대적인 상호 작용을 통해 공동으로 개선되는 공진화 강화 학습 프레임워크인 EvoRubrics를 제안합니다. 정책이 평가 기준 생성기의 지침에 따라 개선됨에 따라, 평가 기준 생성기는 그 기준을 조정하여 구별력과 정보성을 유지하고, 이를 통해 평가는 정책을 실시간으로 추적할 수 있으며 자연스럽게 자동적인 교육 과정을 유도합니다. 실험 결과는 EvoRubrics가 다양한 벤치마크에서 정적 및 동적 평가 기준 기반 모델보다 일관되게 우수한 성능을 보임을 보여줍니다. 학습된 평가 기준 생성기는 전송 가능한 보상 모델로 추가적으로 활용될 수 있습니다. 주목할 점은 외부 감독 없이도 완전히 자체 지도 방식으로 작동하는 변형이 의미 있는 성능 향상을 달성한다는 것입니다. 이는 생성과 평가 간의 공진화만으로도 충분히 풍부한 학습 신호를 제공할 수 있음을 시사합니다. 저희 코드는 다음 링크에서 공개적으로 이용하실 수 있습니다: https://anonymous.4open.science/r/EvoRubrics-2155/.
Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre-constructed rubrics remain static throughout training, creating a fundamental mismatch with the evolving policy: fixed criteria gradually lose discriminative power as the model improves, leading to reward saturation and potential hacking. Recent dynamic rubric methods partially address this but rely on external frontier models or ground-truth answers, and update rubrics only at coarse granularity. We propose EvoRubrics, a co-evolutionary RL framework where a Policy LLM and a Rubric Generator jointly improve through adversarial interaction within each training step. As the policy improves under the rubric generator's guidance, the rubric generator adapts its criteria to remain discriminative and informative, enabling evaluation to track the policy in real time and naturally inducing an automatic curriculum. Experiments show that EvoRubrics consistently outperforms static and dynamic rubric baselines across benchmarks. The learned Rubric Generator further generalizes as a transferable reward model. Notably, even a fully self-supervised variant without any external supervision achieves meaningful gains, suggesting that co-evolution between generation and evaluation alone can provide sufficiently rich learning signals. Our code is publicly available at https://anonymous.4open.science/r/EvoRubrics-2155/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.