2607.07436v1 Jul 08, 2026 cs.AI

맹인 큐레이터: 편향된 평가자가 어떻게 자체 진화 에이전트의 기술 퇴화를 은밀하게 방해하는가

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

Ya Cui
Ya Cui
Citations: 43
h-index: 4
Guanghui Wang
Guanghui Wang
Citations: 23
h-index: 3
Pei-Gen He
Pei-Gen He
Citations: 21
h-index: 3
Wei Qiu
Wei Qiu
School of Computer Science and Engineering, Nanyang Technological University, Singapore
Citations: 498
h-index: 8
Xing Zhang
Xing Zhang
Citations: 164
h-index: 9
Ziyuan Li
Ziyuan Li
Citations: 104
h-index: 4
Bing Zhu
Bing Zhu
Citations: 39
h-index: 4

자체 진화 에이전트는 실패를 관찰하여 불필요한 기술을 제거합니다. 그렇다면, 평가자가 이러한 실패를 감지할 수 없을 때 어떤 일이 발생할까요? 기술 퇴지는 성장하는 기술 목록이 최소 수준 이하로 떨어지지 않도록 하는 중요한 제약 조건이지만, 이는 편향되지 않은 보상을 전제로 합니다. 그러나 레퍼런스 없이 수행되는 작업에서 사용되는 LLM 평가자는 이러한 가정을 위배합니다. 우리는 편향된 평가자가 단순히 노이즈를 추가하는 것이 아니라, extit{큐레이터 기능을 은밀하게 비활성화}한다는 것을 보여줍니다. 우리는 부패한 보상 분석을 통해 이를 명확히 하고, 결정적인 보상에 인위적으로 오류를 주입하여 원인과 결과 관계를 규명했습니다. 또한, 코드 생성 검증을 포함하는 레퍼런스 없는 보고서 작성 테스트 환경에서 행동 연구를 수행했습니다. 대칭적인 노이즈는 기술 퇴지에 영향을 미치지 않지만, extit{오탐(실패가 성공으로 인식되는 현상)} 편향은 특정 임계값을 초과하면 기여도 기반의 기술 퇴화를 완전히 무력화하며, 아무리 많은 데이터를 추가해도 이 임계값을 넘을 수 없습니다. 진정한 기술 퇴자와 성능 제한으로 인한 변화를 구분함으로써, 이러한 extit{메커니즘적 결함}이 보편적으로 발생하며, 도메인과 실패율에 관계없이 거의 오탐이 없는 검증기와 유사한 평가자에게만 영향을 미치지 않는다는 것을 확인했습니다. 하지만 결과는 환경에 따라 다릅니다. 평가 품질은 동일한 오류가 기술 융합을 방해하는 경우에만 저하되며, 다른 경우에는 안정적으로 유지됩니다. 따라서 비활성화된 큐레이터는 extit{눈에 띄지 않게 작용하며}, 어떠한 집계 지표에도 반영되지 않습니다. 본 연구의 기여는 성능 향상이 아니라 안전성 확보에 있습니다. 간단한 오류 주입 감사 절차를 통해 운영자는 배포 전에 평가자가 어느 쪽 임계값에 속하는지 확인할 수 있습니다.

Original Abstract

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks force upon us. We show that a biased judge does not merely add noise; it \emph{silently switches off the curator}. We make this precise with a corrupted-reward analysis and, isolating the causal channel by injecting corruption on top of a deterministic reward, a behavioral study on a reference-free report-writing testbed with a code-generation cross-check. Symmetric noise leaves retirement intact, but \emph{false-pass} bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this \emph{mechanism} failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream \emph{outcome}, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is \emph{silent}, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!