ScholarQuest: 체계 분류 기반의 에이전트 기반 학술 논문 검색 벤치마크 (개방형 문헌 환경)
ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments
학술 논문 검색은 과학 연구의 핵심 단계이며, LLM(대규모 언어 모델) 기반 검색 에이전트는 반복적이고 목적 지향적인 문헌 탐색을 위한 유망한 패러다임으로 부상하고 있습니다. 그러나 기존 벤치마크는 실제 개방형 문헌 환경에서 에이전트 기반 학술 논문 검색을 체계적으로 평가하기에 충분하지 않습니다. 본 연구에서는 대규모의 체계 분류 기반 에이전트 기반 학술 논문 검색 벤치마크인 ScholarQuest를 제안합니다. ScholarQuest는 1,000개 이상의 컴퓨터 과학 주제와 방법론 중심, 상황 기반, 비교 기반, 범위 제한 등 네 가지 대표적인 연구 의도를 포함하여 구축되었습니다. 또한 확장 가능한 답변 생성 기능을 제공하며, 재현 가능한 평가를 위한 공유 검색 백엔드인 ScholarBase를 함께 제공합니다. 벤치마크 결과는 에이전트 기반 방법이 단일 단계 검색 기준보다 우수한 성능을 보이지만, 최고 성능을 보이는 에이전트조차 Recall@100에서 0.314, Recall@All에서 0.355의 낮은 수치를 기록하여, 개선될 여지가 매우 많음을 보여줍니다. 또한 검색 효율성, 의도 수준별 안정성 및 실패 사례에 대한 분석은 ScholarQuest가 학술 논문 검색 에이전트에 대한 다차원적인 평가 지표를 제공하는 데 효과적임을 강조합니다.
Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative, intent-driven literature exploration. However, existing benchmarks are insufficient for systematically evaluating agentic academic search under realistic open literature environments. We propose ScholarQuest, a large-scale, taxonomy-guided benchmark for agentic academic paper search. ScholarQuest is constructed from over 1,000 computer science topics and four representative research intents, including method-oriented, setting-anchored, comparison-based, and scope-controlled queries. It further provides scalable answer construction and a shared retrieval backend ScholarBase for reproducible evaluation. Benchmarking results show that agentic methods outperform single-shot retrieval baselines, yet the best-performing agent only achieves 0.314 Recall@100 and 0.355 Recall@All, indicating substantial room for improvement. In addition, analyses of search efficiency, intent-level robustness, and failure cases further highlight the benchmark's ability to provide multi-dimensional evaluation signals for academic paper search agents.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.