2607.24300v1 Jul 27, 2026 cs.CL

휴리스틱 기반 자기 개선 에이전트에서 자체 검증은 신뢰성이 낮음

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

Cong Cao
Cong Cao
Citations: 25
h-index: 3
Fangfang Yuan
Fangfang Yuan
Citations: 244
h-index: 5
Diandian Guo
Diandian Guo
Citations: 16
h-index: 2
Yingqi Wang
Yingqi Wang
Citations: 0
h-index: 0
Yueshan Wang
Yueshan Wang
Citations: 39
h-index: 3
Dakui Wang
Dakui Wang
Citations: 9
h-index: 2

자기 개선 에이전트는 절차적 정책, 제어기 또는 휴리스틱 규칙을 반복적으로 수정하면서 능력을 향상시킵니다. 이들은 일반적으로 이후 변경 사항을 수용할지 결정하기 위해 스스로 작성한 테스트나 지표에 의존합니다. 에이전트는 최적화 대상과 검증 도구를 모두 통제하므로, 자체 할당된 점수가 거의 완벽하게 유지되는 반면 실제 배포 성능은 저하되거나 낮게 유지될 수 있습니다. 본 연구에서는 검증-배포 격차라는 현상을 통해 이 문제를 분석합니다. 이 격차는 에이전트가 관찰하거나 접근할 수 없는 외부 평가와 에이전트가 자체적으로 작성한 검증 신호 간의 불일치를 의미합니다. 우리는 반복적인 정책 및 테스트 재작성 과정에서 자체 검증이 어떻게 실패하는지, 이러한 실패가 능력 수준에 따라 어떻게 변화하는지, 그리고 실제 성능 저하를 방지하기 위해 얼마나 적은 외부 신뢰가 필요한지를 연구합니다. 이 문제를 해결하기 위해 Sealed Exogenous Acceptance Loop (SEAL)을 제안합니다. SEAL은 자체 작성된 테스트를 유지하면서, 고정된 하드웨어 기반 감사 시스템을 통해 각 후보 솔루션을 기존 솔루션과 비교합니다. 에이전트는 감사를 작성하거나 검사할 수 없으며, 단순히 수락/거부 여부만 받고, 명확한 성능 저하가 발생하면 전체 기존 상태를 유지합니다. 실험 결과는 이 문제가 휴리스틱 학습 환경에서 자주 발생하는 것을 보여줍니다. 이러한 환경은 대상 목표의 시행착오 기반 발견을 필요로 합니다. 또한 자체 작성된 검증 실패는 능력 수준에 따라 계층화되는 경향이 있음을 확인했습니다. 약한 에이전트는 쉬운 자체 테스트 뒤에 이미 획득한 전략을 손상시키는 반면, 더 강력한 에이전트는 더 안정적이지만 여전히 배포 환경의 분포를 부정확하게 측정합니다. 일반적인 자체 작성된 제약 조건은 이 격차를 신뢰성 있게 해소하지 못합니다. 이에 비해 SEAL은 여섯 가지 모델과 세 개의 랜덤 시드에서 보호되지 않은 기본 설정보다 우수한 성능을 보였습니다. 신뢰할 수 있는 자기 개선은 자체 검증을 완전히 포기할 필요는 없지만, 에이전트의 통제 범위를 벗어난 최소한 하나의 배포 승인 신호가 필요합니다.

Original Abstract

Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!