2606.29713v1 Jun 29, 2026 cs.CL

SEVA: 프로세스 보상을 활용한 자기 진화 검증 에이전트

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

Yi Nian
Yi Nian
Citations: 141
h-index: 6
Haiyu Zhang
Haiyu Zhang
Citations: 14
h-index: 2
Yue Zhao
Yue Zhao
Citations: 193
h-index: 7
Aojie Yuan
Aojie Yuan
Citations: 11
h-index: 2
Zijian Su
Zijian Su
Citations: 0
h-index: 0

LLM 기반 에이전트의 신뢰성 저하의 주된 원인은 '환각' 현상이며, 사실 출처 검증 도구는 이를 해결하기 위한 최후의 방어선입니다. 그러나 현재의 검증 도구들은 불투명한 이진 레이블만 제공하여 에이전트가 스스로 오류를 수정하거나 운영자가 감사를 수행하는 것을 어렵게 만듭니다. 본 논문에서는 증거 정렬, 단계별 추론 과정, 교정된 신뢰도, 그리고 실행 가능한 해결책을 제시하는 6가지 범주의 오류 진단 정보를 제공하는 구조화된 검증 에이전트인 SEVA를 소개합니다. 이러한 에이전트를 강화 학습으로 훈련하는 것은 쉽지 않습니다. 다중 구성 요소 출력에 대한 표준 이진 보상은 '장점 붕괴' 현상을 유발하며, 이는 그룹 내 보상 편차를 감소시키고 GRPO 그래디언트 소실을 초래합니다. 우리는 이러한 문제를 프로세스 보상을 통해 해결했습니다. 프로세스 보상은 검증 품질을 5개의 독립적인 구성 요소로 분해하고, 이들을 70/30의 가중치를 부여하여 프로세스 신호에 더 큰 비중을 두어 그래디언트를 복원하고 암묵적인 학습 과정을 유도합니다. 에이전트는 먼저 검증 동작(정렬 정확도 0.917 -> 0.997, 형식 준수율 72% -> 100%)을 익힌 후 결과 도출 능력을 향상시킵니다 (F1 점수 64.9 -> 69.0). 구조화된 출력은 '검증 -> 성찰 -> 탐색 -> 개선'의 자기 진화 루프를 가능하게 하며, 7B 모델에 대해 4번 반복했을 때 예상치 못한 구조적 발견을 이끌어냈습니다. 각 단계에서 에이전트는 일반적인 능력을 갖추는 것이 아니라 특정 벤치마크에 특화된 능력을 갖추게 됩니다 (+15 pp on HaluEval, 동일 모델에서 TruthfulQA의 경우 -10 ~ -14 pp). ClearFacts 데이터셋에서는 SEVA-3B가 GPT-4o-mini와 유사한 성능(F1 점수 69.0 vs. 69.8)을 보이지만, 훨씬 풍부하고 감사 가능한 출력을 제공합니다. 이는 다음과 같은 원칙을 확인시켜줍니다. 다중 구성 요소 생성을 포함하는 모든 강화 학습 작업에서 보상의 세분성은 출력의 세분성과 일치해야 합니다.

Original Abstract

Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers emit only opaque binary labels, leaving agents unable to self-correct and operators unable to audit. We present SEVA, a structured verification agent that emits evidence alignments, step-by-step reasoning chains, calibrated confidence, and a six-category error diagnosis with actionable fixes. Training such an agent with RL is non-trivial: standard binary reward on multi-component output triggers advantage collapse -- within-group reward variance vanishes and the GRPO gradient disappears. We resolve this with a process reward that decomposes verification quality into five independent components weighted 70/30 toward process signals, restoring the gradient and inducing an implicit curriculum -- the agent first masters verification behavior (alignment 0.917 -> 0.997, format 72% -> 100%), then outcomes (F1 64.9 -> 69.0). Structured output further enables a Verify -> Reflect -> Probe -> Refine self-evolution loop, which over four rounds on a 7B model surfaces an unexpected structural finding: each round produces a benchmark-specialist, not a generalist (+15 pp on HaluEval, -10 to -14 pp on TruthfulQA in the same model, persistent at 4x data). On ClearFacts, SEVA-3B matches GPT-4o-mini (69.0 vs. 69.8 F1) while producing substantially richer, auditable output -- confirming a principle that should generalize: for any RL task with multi-component generation, reward granularity must match output granularity.

1 Citations
0 Influential
3.5 Altmetric
18.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!