2606.13262v1 Jun 11, 2026 cs.AI

판결에서 프로세스로: 다단계 사실 검증을 위한 에이전트 기반 강화 학습

From Verdict to Process: Agentic Reinforcement Learning for Multi-Stage Fact Verification

Shenghong He
Shenghong He
Citations: 0
h-index: 0
Siyu Zhu
Siyu Zhu
Citations: 202
h-index: 4
Chao Yu
Chao Yu
Citations: 1
h-index: 1
Rongxin Yang
Rongxin Yang
Citations: 8
h-index: 2

최근 대규모 언어 모델(LLM)과 검색 증강 추론을 결합한 접근 방식은 자동화된 사실 검증에 유망한 결과를 보여주었습니다. 복잡한 주장을 처리하기 위해 이러한 검증 파이프라인은 일반적으로 주장 분해, 증거 수집 및 판결 예측을 포함하는 긴밀하게 결합된 모듈을 조정하는 다단계 워크플로우를 실행합니다. 그러나 기존 방법은 개별 단계를 독립적으로 최적화하거나 고정된 휴리스틱에 의존하며, 이는 단계 간의 적응적인 조정을 제한하고 최적이 아닌 결과를 초래할 수 있습니다. 본 연구에서는 다단계 사실 검증 경로의 엔드 투 엔드 최적화를 위한 에이전트 기반 강화 학습 프레임워크인 ProFact를 제안합니다. ProFact는 통합 정책을 훈련하여 주장 분해, 증거 탐색, 답변 생성 및 판결 예측을 조정합니다. 최종 진실성 레이블에서 제공되는 희소하고 지연된 감독 문제를 해결하기 위해, ProFact는 검증 프로세스 전반에 걸쳐 단계 수준의 학습 신호를 제공하는 프로세스 인식 보상을 도입합니다. 실험적 평가는 ProFact가 검증 성능과 추론 효율성 모두에서 강력한 기준 모델보다 지속적으로 우수한 성능을 발휘한다는 것을 보여줍니다. 이러한 결과는 다단계 사실 검증을 위한 프로세스 인식 경로 최적화의 효과를 강조합니다.

Original Abstract

Recent approaches combining Large Language Models (LLMs) with retrieval-augmented reasoning have shown promise for automated fact verification. To process complex claims, these verification pipelines typically execute multi-stage workflows that coordinate tightly coupled modules, including claim decomposition, evidence gathering, and verdict prediction. However, existing methods optimize individual stages in isolation or rely on fixed heuristics, which limits adaptive coordination among stages and can lead to suboptimal outcomes. In this work, we propose ProFact, an agentic reinforcement learning framework for end-to-end optimization of multi-stage fact verification trajectories. ProFact trains a unified policy to coordinate claim decomposition, evidence seeking, answer generation, and verdict prediction. To address the sparse and delayed supervision provided by final veracity labels, ProFact introduces process-aware rewards that provide stage-level learning signals throughout the verification process. Empirical evaluation shows that ProFact consistently outperforms strong baselines in both verification performance and inference efficiency. These results highlight the effectiveness of process-aware trajectory optimization for multi-stage fact verification.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!