2602.07549v1 Feb 07, 2026 cs.AI

언제 충분한 것이 충분하지 않은가? 검색 에이전트의 완료 착각

When Is Enough Not Enough? Illusory Completion in Search Agents

Moontae Lee
Moontae Lee
Citations: 37
h-index: 5
Dayoon Ko
Dayoon Ko
Citations: 47
h-index: 4
Jihyuk Kim
Jihyuk Kim
Citations: 8
h-index: 1
Sohyeon Kim
Sohyeon Kim
Seoul National University
Citations: 9
h-index: 1
Haeju Park
Haeju Park
Citations: 19
h-index: 2
Dahyun Lee
Dahyun Lee
Citations: 19
h-index: 2
Gunhee Kim
Gunhee Kim
Citations: 134
h-index: 7
Kyungjae Lee
Kyungjae Lee
Citations: 114
h-index: 5

최근 검색 에이전트들은 멀티턴 추론과 검색 도구를 활용하여 멀티홉 및 장기 수행 벤치마크에서 강력한 성능을 달성하고 있습니다. 그러나 에이전트가 질문에 내재된 여러 조건들을 추적, 검증, 유지함으로써 모든 요구사항에 대해 신뢰할 수 있는 추론을 수행하는지는 여전히 불분명합니다. 본 연구에서는 유효한 답변이 여러 제약 조건을 동시에 만족해야 하는 다중 제약 문제 상황에서 이러한 능력을 분석합니다. 연구 결과, 제약 조건이 해결되지 않았거나 위배되었음에도 불구하고 에이전트가 작업이 완료되었다고 믿는 '완료 착각(Illusory Completion)' 현상이 빈번하게 발생하며, 이것이 검증이 부족한 답변으로 이어진다는 것을 발견했습니다. 이러한 행동을 진단하기 위해, 우리는 멀티턴 추론 전반에 걸쳐 각 제약 조건에 대한 증거적 지지와 에이전트의 믿음을 추적하는 평가 프레임워크인 'Epistemic Ledger'를 제안합니다. 우리의 분석은 네 가지 반복적인 실패 패턴인 근거 없는 단언(bare assertions), 반박 간과(overlooked refutations), 정체(stagnation), 조기 종료(premature exit)를 밝혀냈습니다. 이러한 발견에 기반하여, 우리는 추론 시점 추적기인 'LiveLedger'를 통해 실행 중 명시적인 제약 상태 추적이 이러한 실패를 완화할 수 있는지 조사합니다. 이 간단한 개입은 다중 제약 문제에서 성능을 일관되게 향상시켰으며, 검증 부족 답변을 대폭(최대 26.5%) 줄이고 전반적인 정확도를(최대 11.6%) 개선하는 성과를 보였습니다.

Original Abstract

Recent search agents leverage multi-turn reasoning and search tools to achieve strong performance on multi-hop and long-horizon benchmarks. Yet it remains unclear whether they reliably reason across all requirements by tracking, verifying, and maintaining multiple conditions in these questions. We study this capability under multi-constraint problems, where valid answers must satisfy several constraints simultaneously. We find that illusory completion frequently occurs, wherein agents believe tasks are complete despite unresolved or violated constraints, leading to underverified answers. To diagnose this behavior, we introduce the Epistemic Ledger, an evaluation framework that tracks evidential support and agents' beliefs for each constraint throughout multi-turn reasoning. Our analysis reveals four recurring failure patterns: bare assertions, overlooked refutations, stagnation, and premature exit. Motivated by these findings, we examine whether explicit constraint-state tracking during execution mitigates these failures via LiveLedger, an inference-time tracker. This simple intervention consistently improves performance, substantially reducing underverified answers (by up to 26.5%) and improving overall accuracy (by up to 11.6%) on multi-constraint problems.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!