2608.01913v1 Aug 03, 2026 cs.AI

장기 탐색 에이전트의 검색 행동 및 실패 요인 진단

Diagnosing Search Behavior and Failure Modes in Long-Horizon Search Agents

Jiaxin Mao
Jiaxin Mao
Citations: 159
h-index: 7
Fengbin Zhu
Fengbin Zhu
National University of Singapore
Citations: 1,375
h-index: 14
Tat-Seng Chua
Tat-Seng Chua
Citations: 2,451
h-index: 22
Qi Liu
Qi Liu
Citations: 145
h-index: 6

심층 검색 에이전트는 어려운 정보 검색 질문에 대한 답변을 얻기 위해 반복적으로 검색 쿼리를 실행하여 관련 증거를 수집하지만, 더 많은 검색 노력이 실제로 더 나은 답변으로 이어지는지 여부는 아직 명확하지 않습니다. 본 연구는 장기 탐색 에이전트의 행동 경로 수준 분석을 통해 이러한 질문에 대한 답을 찾고자 합니다. 인간이 직접 평가한 문서 수준의 관련성 판단 데이터를 사용하여 각 검색 단계에서 수집된 증거를 평가하고, 에이전트의 행동을 두 가지 측면으로 구분합니다. 즉, 에이전트가 어떤 증거를 수집하는지, 그리고 해당 증거를 얼마나 효과적으로 사용하는지를 분석합니다. 이러한 구분을 통해 실패 원인을 증거 회수 부족(필요한 증거가 전혀 발견되지 않는 경우)과 활용 부족(관련된 증거가 수집되었지만 제대로 사용되지 않는 경우)으로 세분화할 수 있습니다. 본 연구에서는 검색 모델 및 평가 시스템을 고정한 상태에서 BrowseComp-Plus 데이터셋을 사용하여 6개의 에이전트를 비교하고, open-web 검색 API를 통해 BrowseComp 데이터셋에서도 동일한 결과를 검증합니다. 다양한 환경에서 실험 결과, 검색 노력과 답변 품질은 약하게 연관되어 있음을 확인했습니다. 답변 정확도는 검색 횟수나 소비된 컨텍스트 양보다 수집된 증거의 품질, 특히 누적 검색 재현율과 더 밀접하게 관련되어 있습니다. 유용한 증거는 종종 탐색 경로 초기에 나타나는 반면, 에이전트는 계속해서 검색을 수행하여 낮은 효율성을 보이는 쿼리들이 연달아 발생하는 경향을 보입니다. 쿼리 수준에서, 탐색적인 재정의는 여전히 유용하지만, 가장 성능이 좋은 에이전트는 훨씬 적은 수의 중복 쿼리를 실행합니다. 종합적으로, 본 연구는 장기 탐색 에이전트의 검색 행동 및 실패 요인을 체계적으로 분석함으로써, 더 강력한 쿼리 생성, 효과적인 증거 선택 및 컨텍스트 관리, 그리고 충분한 관련 증거가 수집되었는지 여부를 기준으로 하는 중단 기준을 포함하는, 향상된 심층 연구 시스템 개발을 위한 실질적인 방향을 제시합니다.

Original Abstract

Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnosis of long-horizon search agents. Using human-annotated document-level relevance judgments, we evaluate the evidence retrieved at each search step and separate two stages of agent behavior: what evidence an agent retrieves and how effectively it uses that evidence. This distinction further allows us to decompose failures into retrieval gaps, where the necessary evidence is never found, and utilization gaps, where relevant evidence is retrieved but not used correctly. With the retrieval model and evaluation harness held fixed, we compare six agents on BrowseComp-Plus and further validate our findings on BrowseComp with an open-web search API. Across settings, we find that search effort and answer quality are only weakly aligned. Answer accuracy is better correlated with the quality of retrieved evidence, especially cumulative retrieval recall, than with the number of searches or the amount of context consumed. Useful evidence often appears early in the trajectory, yet agents tend to continue searching, producing a long tail of low-yield retrieval steps. At the query level, exploratory reformulations remain useful, but the best-performing agents issue far fewer redundant queries. Overall, by systematically characterizing the search behavior and failure modes of long-horizon search agents, this work points to practical directions for building better deep research systems, including stronger query formulation, more effective evidence selection and context management, and stopping criteria based on whether sufficient supporting evidence has been retrieved.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!