2606.23038v1 Jun 22, 2026 cs.LG

EvoRubrics: 적대적 공진화를 통한 LLM 강화 학습을 위한 동적인 평가 기준

EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning

Junfeng Zhao
Junfeng Zhao
Citations: 576
h-index: 12
Yasha Wang
Yasha Wang
Citations: 1,053
h-index: 15
Baixiang Huang
Baixiang Huang
Citations: 374
h-index: 10
Yue Fang
Yue Fang
Citations: 117
h-index: 6
Hongxin Ding
Hongxin Ding
Citations: 100
h-index: 6
Weibin Liao
Weibin Liao
Citations: 23
h-index: 3
Zheng Li
Zheng Li
Citations: 49
h-index: 2
Jinyang Zhang
Jinyang Zhang
Citations: 10
h-index: 2
Zhijing Wu
Zhijing Wu
Citations: 116
h-index: 7

평가 기준 기반 보상은 검증 가능한 답변이 없는 개방형 작업에서 강화 학습을 위한 해석 가능하고 세밀한 최적화 신호를 제공합니다. 그러나 미리 구성된 평가 기준은 훈련 과정 전반에 걸쳐 고정되어 있으므로, 정책의 변화와 근본적인 불일치를 야기합니다. 즉, 모델이 개선됨에 따라 고정된 기준은 점차 구별력을 잃게 되어 보상 포화 및 잠재적인 해킹으로 이어질 수 있습니다. 최근 동적 평가 기준 방법은 이러한 문제를 부분적으로 해결하지만, 외부 최첨단 모델이나 정답을 사용하고, 평가 기준을 조잡한 수준에서만 업데이트합니다. 우리는 정책 LLM과 평가 기준 생성기가 각 훈련 단계 내에서 적대적인 상호 작용을 통해 공동으로 개선되는 공진화 강화 학습 프레임워크인 EvoRubrics를 제안합니다. 정책이 평가 기준 생성기의 지침에 따라 개선됨에 따라, 평가 기준 생성기는 그 기준을 조정하여 구별력과 정보성을 유지하고, 이를 통해 평가는 정책을 실시간으로 추적할 수 있으며 자연스럽게 자동적인 교육 과정을 유도합니다. 실험 결과는 EvoRubrics가 다양한 벤치마크에서 정적 및 동적 평가 기준 기반 모델보다 일관되게 우수한 성능을 보임을 보여줍니다. 학습된 평가 기준 생성기는 전송 가능한 보상 모델로 추가적으로 활용될 수 있습니다. 주목할 점은 외부 감독 없이도 완전히 자체 지도 방식으로 작동하는 변형이 의미 있는 성능 향상을 달성한다는 것입니다. 이는 생성과 평가 간의 공진화만으로도 충분히 풍부한 학습 신호를 제공할 수 있음을 시사합니다. 저희 코드는 다음 링크에서 공개적으로 이용하실 수 있습니다: https://anonymous.4open.science/r/EvoRubrics-2155/.

Original Abstract

Rubric-based rewards offer interpretable and fine-grained optimization signals for reinforcement learning in open-ended tasks where verifiable answers are unavailable. However, pre-constructed rubrics remain static throughout training, creating a fundamental mismatch with the evolving policy: fixed criteria gradually lose discriminative power as the model improves, leading to reward saturation and potential hacking. Recent dynamic rubric methods partially address this but rely on external frontier models or ground-truth answers, and update rubrics only at coarse granularity. We propose EvoRubrics, a co-evolutionary RL framework where a Policy LLM and a Rubric Generator jointly improve through adversarial interaction within each training step. As the policy improves under the rubric generator's guidance, the rubric generator adapts its criteria to remain discriminative and informative, enabling evaluation to track the policy in real time and naturally inducing an automatic curriculum. Experiments show that EvoRubrics consistently outperforms static and dynamic rubric baselines across benchmarks. The learned Rubric Generator further generalizes as a transferable reward model. Notably, even a fully self-supervised variant without any external supervision achieves meaningful gains, suggesting that co-evolution between generation and evaluation alone can provide sufficiently rich learning signals. Our code is publicly available at https://anonymous.4open.science/r/EvoRubrics-2155/.

2 Citations
0 Influential
7.5 Altmetric
39.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!