2608.03527v1 Aug 04, 2026 cs.IR

심층 연구 에이전트를 위한 검색 기준 기반 문서 재순위화 방법

Training Documents Reranker with Search Rubrics for Deep Research Agent

Yutao Zhu
Yutao Zhu
University of Montreal
Citations: 4,882
h-index: 29
Zhicheng Dou
Zhicheng Dou
Citations: 3,020
h-index: 28
Hao Wang
Hao Wang
Citations: 242
h-index: 8
Haijin Liang
Haijin Liang
Citations: 35
h-index: 4
Tong Zhao
Tong Zhao
Citations: 31
h-index: 4
Haibo Shi
Haibo Shi
Citations: 201
h-index: 6
Yu Lu
Yu Lu
Citations: 87
h-index: 5
Qiaolin Xia
Qiaolin Xia
Citations: 515
h-index: 9
Jianfei Xi
Jianfei Xi
Citations: 0
h-index: 0
Wenhan Liu
Wenhan Liu
Citations: 673
h-index: 7
Hui-Qing Xu
Hui-Qing Xu
Citations: 2
h-index: 1

검색 시스템은 관련 문서를 제공하여 심층 연구 에이전트가 고품질 답변을 생성하도록 돕습니다. 그러나 기존의 검색 시스템은 일반적으로 관련성 매칭을 통해 문서를 선택합니다. 하지만 개별적으로는 잘 맞는 상위 $k$개의 문서들이, 에이전트 질의의 복잡한 정보 요구 사항(예: 다양성, 간결성 및 권위)을 충족하는 extit{집합}을 형성하지 못할 수 있습니다. 본 논문에서는 고품질 문서 집합이 각 에이전트 질의에 대해 충족해야 하는 요건을 명시적으로 정의하는 검색 기준을 제안합니다. 이러한 검색 기준은 계층적 구조로 구성되어 있으며, 강력한 LLM을 사용하여 합성됩니다. 제안된 검색 기준을 기반으로, 우리는 검색된 문서 중에서 고품질 부분집합을 선택하기 위한 문서 재순위화 모델인 extbf{RubricRanker}를 학습시킵니다. 우리는 두 단계의 학습 프레임워크를 설계했으며, 이는 기준 지향적 지도 학습과 기준 기반 강화 학습으로 구성됩니다. 광범위한 실험 결과는 RubricRanker가 네 가지 심층 연구 벤치마크에서 가장 강력한 기본 모델보다 2.6점 더 높은 성능을 보이며, 다섯 개의 RAG 벤치마크에서도 우수한 일반화 능력을 보여준다는 것을 입증합니다.

Original Abstract

Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!