2607.20083v1 Jul 22, 2026 cs.LG

DynamicRubric을 이용한 LLM 평가 모델 및 정책의 공진화

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Qingyao Ai
Qingyao Ai
Citations: 1,764
h-index: 22
Weihang Su
Weihang Su
Citations: 797
h-index: 18
Min Zhang
Min Zhang
Citations: 104
h-index: 6
Beining Wang
Beining Wang
Citations: 8
h-index: 1
Hongtao Tian
Hongtao Tian
Citations: 42
h-index: 4
Hao Kong
Hao Kong
Citations: 0
h-index: 0
Tao Yang
Tao Yang
Citations: 20
h-index: 2
Ting Yao
Ting Yao
Citations: 18
h-index: 2
Qingyi Pan
Qingyi Pan
Citations: 11
h-index: 2
Yueyue Wu
Yueyue Wu
Citations: 319
h-index: 10
Yiqun Liu
Yiqun Liu
Citations: 25
h-index: 3

대규모 언어 모델(LLM)을 개선하는 주요 방법 중 하나는, 정책에 의해 생성된 샘플 데이터에 대한 평가 모델 피드백을 활용하여 추가 학습을 수행하는 것입니다. 정책이 발전함에 따라, 이러한 샘플 응답들은 품질 면에서 점점 더 유사해집니다. 이와 같은 유사성은 정책 최적화의 병목 현상을 야기합니다. 응답 간 상대적인 평가 점수 차이가 줄어들면, 이는 약하거나 오해를 불러일으킬 수 있는 정책 감독 신호로 작용합니다. 본 연구에서는 확률 할당 관점에서 이러한 점수 차이가 중요한 이유를 이론적으로 분석했습니다. 분석 결과, 한 응답에서 다른 응답으로 확률 질량이 이동하는 방향상의 이득은 두 응답 간의 상대적인 평가 점수 차이와 정확히 일치한다는 것을 보여줍니다. 이는 상대적인 점수 차이를 정책 최적화를 위한 핵심 신호로 식별합니다. 이러한 관점에서, 본 연구에서는 DynamicRubric이라는 응답 집합 기반의 평가 모델-정책 공진화 프레임워크를 제안합니다. DynamicRubric은 각 후보 집합에 대해 가중치가 부여된 이분법적인 평가 항목을 생성하고, 이를 통해 얻어진 판단들을 통합하여 응답 수준의 점수를 산출합니다. 80억 개의 파라미터를 가진 모델을 대상으로 실험한 결과, DynamicRubric은 기존 방식보다 평가 모델 성능을 향상시키고 더욱 강력한 정책 감독 신호를 제공했습니다. 또한, DynamicRubric으로 최적화된 정책은 검증 가능한 추론 및 코딩 작업에서 더 나은 성능을 보였습니다. DynamicRubric으로 최적화된 모델은 WeChat Search의 AI 답변 기능에 완전히 적용되었으며, 매일 수백만 건의 요청을 처리하는 온라인 트래픽 전체를 지원하며 주요 온라인 지표를 개선했습니다. 이러한 결과는 평가 모델 기반의 추가 학습에 대한 중요한 원칙을 제시합니다: 평가 모델은 감독하는 정책과 함께 발전해야 합니다.

Original Abstract

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a bottleneck for policy optimization: collapsed relative evaluator score gaps yield weak or misleading policy supervision. We theoretically characterize why these gaps matter through a probability allocation view, showing that the directional gain of shifting probability mass from one response to another is exactly the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. Motivated by this view, we propose DynamicRubric, a response-set-conditioned evaluator--policy co-evolution framework that generates weighted binary rubric items for each candidate set and aggregates the resulting judgments into response-level scores. In our experiments with 8B backbones, DynamicRubric improves evaluator performance and provides stronger policy supervision than baselines using a 70B reward model or a 235B static rubric generator. DynamicRubric-optimized policies also show gains on verifiable reasoning and coding tasks. A DynamicRubric-optimized model is fully deployed in WeChat Search's AI answering scenario, where it serves all online traffic across tens of millions of requests per day and improves key online metrics. These results suggest a principle for evaluator-guided post-training: evaluators should evolve with the policies they supervise.

1 Citations
0 Influential
11 Altmetric
56.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!