2607.23524v1 Jul 26, 2026 cs.AI

심층 검색에서의 위임 지능: 분리된 능력 진단을 위한 제어 가능한 프레임워크

Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis

Xinhao Yao
Xinhao Yao
Citations: 54
h-index: 5
Yuyao Zhang
Yuyao Zhang
Citations: 676
h-index: 6
Ruifeng Ren
Ruifeng Ren
Citations: 58
h-index: 5
Haoran Tan
Haoran Tan
Citations: 205
h-index: 4
Yuanzhuo Liu
Yuanzhuo Liu
Citations: 0
h-index: 0
Changhao Wang
Changhao Wang
Citations: 0
h-index: 0
Yunfei Yu
Yunfei Yu
Citations: 0
h-index: 0
Minlong Peng
Minlong Peng
Citations: 46
h-index: 4
Yong Liu
Yong Liu
Citations: 28
h-index: 4

심층 검색은 현대 에이전트 시스템의 핵심 기능으로 자리 잡고 있지만, 일반적으로는 최종 답변 정확도만을 기준으로 평가됩니다. 이러한 통합적인 평가 방식은 정보 검색 품질, 긴 문맥 이해, 증거 검증 및 도구 사용 결정 등을 묶어버려 모델이 실제로 언제 어떻게 정보 검색을 위임해야 하는지 제대로 파악하기 어렵게 만듭니다. 이에 저희는 다음과 같은 목표를 설정했습니다. (1) 심층 검색에서의 이 메타-능력을 '위임 지능'으로 공식화하고, 상호 보완적인 두 가지 차원으로 분해합니다: 검색 의사 결정(정보 부족을 인지하고 언제, 어떻게 검색할지 판단하는 능력) 및 정보 종합 및 검증(다양한 출처에서 증거를 수집하고, 출처의 신뢰성을 평가하며, 노이즈가 많고 잠재적으로 적대적인 환경에서도 정보를 종합하는 능력). (2) 분리되고 재현 가능한 측정을 가능하게 하기 위해 문서 기반 역공학을 활용한 제어 가능한 합성 파이프라인을 개발했습니다. 이를 통해 특정 데이터셋에 국한되지 않고, 제어된 심층 검색 평가를 구축하기 위한 일반적인 방법을 제시합니다. (3) 구체적인 예시로서, 문서 구성 및 도구 접근 방식을 다양하게 변경하여 각 능력 차원을 분리하는 평가 프로토콜과 함께 DelegSearchBench를 구축했습니다. (4) 다양한 모델을 대상으로 실험한 결과, 심층 검색 능력을 최종 답변 정확도만으로 충분히 설명할 수 없다는 것을 입증했습니다.

Original Abstract

Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This coupled evaluation paradigm entangles retrieval quality, long-context comprehension, evidence verification, and tool-use decisions, making it difficult to determine whether a model truly knows when and how to delegate information seeking to search. To this end: (1) We formalize this meta-capability as Delegation Intelligence in deep search and decompose it into complementary dimensions-Search Decision-Making (recognizing information insufficiency and deciding whether, when, and how to search) and Information Synthesis and Verification (aggregating evidence from multiple sources, judging source reliability, and synthesizing information under noisy, potentially adversarial conditions). (2) To enable disentangled and reproducible measurement, we develop a controllable synthesis pipeline built on document-grounded reverse engineering. This yields a general recipe for constructing controlled deep-search evaluations rather than a single fixed dataset. (3) As a concrete instantiation, we construct DelegSearchBench, together with a disentangled evaluation protocol that isolates each capability dimension by varying document composition and tool access. (4) Across representative models, we demonstrate that deep-search competence cannot be adequately characterized by final-answer accuracy alone...

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!