2607.12790v1 Jul 14, 2026 cs.AI

누가 평가자를 평가하는가? 자기 개선 LLM 에이전트를 위한 공진화 평가 지표 및 기술

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

Ya Cui
Ya Cui
Citations: 43
h-index: 4
Guanghui Wang
Guanghui Wang
Citations: 23
h-index: 3
Pei-Gen He
Pei-Gen He
Citations: 21
h-index: 3
Wei Qiu
Wei Qiu
School of Computer Science and Engineering, Nanyang Technological University, Singapore
Citations: 498
h-index: 8
Xing Zhang
Xing Zhang
Citations: 164
h-index: 9
Ziyuan Li
Ziyuan Li
Citations: 104
h-index: 4
Bing Zhu
Bing Zhu
Citations: 39
h-index: 4

자기 진화 에이전트 시스템은 스스로의 기술을 생성, 수정 및 폐기함으로써 성능을 향상시키지만, 이러한 모든 과정은 숨겨진 전제하에 기반합니다. 바로 신뢰할 수 있는 평가 지표가 이미 존재한다는 것입니다. 그러나 실제 응용 분야에서 이는 종종 사실이 아닙니다. 본 논문에서는 세 가지 주장을 제시합니다. 첫째, 평가 지표는 *진화*될 수 있습니다. 저희의 평가 지표 진화 시스템은 작은 결함 감지기의 조합을 사용하여 완전한 진화 라이프사이클을 통해 학습하며, 10개의 기준 항목 집합과 일치하도록 훈련되고, 레이블이 없는 출력에 대한 합의를 통해 정규화되며, 모델이 읽지 못하는 별도의 검증 데이터 세트에 대한 감사 과정을 거쳐 투명하고 검토 가능한 평가 지표를 제공합니다. 둘째, 완벽한 평가 지표가 존재하지 않으므로, 저희는 정확한 평가 지표가 제공했을 경우 달성할 수 있었던 성능을 추정하며, *Double Ratchet*이라는 저희의 공진화 시스템은, 평가 지표와 기술 라이프사이클을 함께 관리함으로써 이를 가능하게 합니다. 코드 생성 (MBPP+), 기업용 텍스트-SQL (Spider~2.0-Snow), 그리고 기준 없이 자유로운 보고서 생성 작업에서, *Double Ratchet* 시스템은 실제 정답 또는 최적의 평가 기준에 의해 구동되는 동일한 기술 루프가 달성한 성능의 88~110%를 유지합니다. 셋째, 안전성은 기준 준수 및 외부 감사를 통해 확보됩니다. 기준 보호 장치를 제거하면 평가 지표는 의미 없는 검출기로 붕괴되지만, 라이프사이클을 제거하는 것은 그렇지 않습니다. 또한, 진화된 기술이 보고서 평가 기준을 악용한 경우, 독립적인 평가자가 이를 발견하고, 하나의 검출기가 이를 수정했으며, 작업에 대한 이해를 가진 평가자는 진화된 출력물을 기존의 기본 라인보다 77%의 쌍에서 더 선호했습니다. 저희는 이러한 오류 발생 가능성을 고려하는 아키텍처가 신뢰할 수 있는 자동 검증 시스템이 존재하지 않는 모든 곳에서 적절한 기본 모델이라고 주장합니다.

Original Abstract

Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We make three claims. First, metrics can be \emph{evolved}: our metric loop searches compositions of small drawback detectors under a full evolutionary lifecycle, trained to agree with a ten-item anchored reference set, regularized by consensus over unlabeled outputs, and audited against a held-out anchor it never reads, yielding a transparent, inspectable metric rather than an opaque judge. Second, since no metric exists to beat, the yardstick is recovering what an accurate metric would have enabled, and \emph{Double Ratchet}, our co-evolution of the metric with a lifecycle-managed skill loop, does so: across code generation (MBPP+), enterprise text-to-SQL (Spider~2.0-Snow), and reference-free report generation, it retains 88--110\% of the held-out lift achieved by the same skill loop driven by ground truth or the best available rubric. Third, safety comes from anchor discipline plus outer audits: removing anchor guards collapses the metric into a vacuous detector while removing the lifecycle does not; and when evolved skills gamed the report rubric, an independent judge caught it, one detector repaired it, and a task-aware judge then preferred the evolved outputs over the pre-evolution baseline in 77\% of decided pairs. We argue this failure-expecting architecture is the right default wherever no reliable automatic verifier exists.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!