문헌 검색 평가 재고찰: 심층 연구는 도움이 되며, 인간이 작성한 참고 문헌 목록은 절대적인 진실이 아니다
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
본 연구에서는 대규모 문헌 검색을 두 가지 상호 보완적인 관점에서 분석합니다. 첫째, 검색 파이프라인 개선에 집중하며, 둘째, 인간의 참고 문헌 목록을 평가 기준으로 활용하는 방법을 검증합니다. 먼저, 전체 논문을 처리하고 인용 네트워크를 기반으로 검색 결과를 확장하는 'Deep Research' 파이프라인을 구현했습니다. 실험 결과, 이 방법은 기존 API만을 사용하는 검색 방식보다 성능이 월등히 우수하며, 250편의 논문으로 구성된 문헌 검색 벤치마크인 RollingEval-Jun25에서 재현율을 20% 미만에서 80% 이상으로 향상시켰습니다. 둘째, 중립적인 LLM(Large Language Model)을 평가 도구로 사용하여 인간의 참고 문헌이 얼마나 정확한 기준인지 조사했습니다. 분석 결과, 인간의 인용 자료 중 51%만이 '어느 정도 관련 있음' 이상의 수준으로 판단되었으며, 이는 최첨단 AI 기반 재순위화 시스템의 86~88%에 비해 현저히 낮은 수치입니다. OpenAlex의 공동 저술 네트워크를 분석한 결과, 인간은 최고의 AI 재순위화 시스템보다 직접적인 협력자를 인용할 가능성이 2.5배 높았습니다. 종합적으로 볼 때, 본 연구는 단일 지표로 문헌 검색을 평가하는 방식에 의문을 제기합니다. 재현율, 주제 관련성 점수, 순위 목록의 다양성, 그리고 공동 저술 거리 분석은 각각 인용 품질의 상호 보완적인 측면을 측정하며, 함께 보고되어야 합니다.
We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the full query paper and expands the retrieved results breadth-first along their bibliographies, and show that it substantially outperforms vanilla API-only search, raising recall on RollingEval-Jun25 (a 250-paper literature-search benchmark) from below 20% to above 80%. Second, we use a neutral LLM-as-a-judge to determine if human references are sound ground truth for the task. We find significant limitations: only 51% of human citations are judged moderately relevant or higher, against 86--88% for the strongest AI-based re-rankers. We study this gap on the OpenAlex co-authorship graph, finding that humans are 2.5x more likely than the best AI re-rankers to cite a direct collaborator. Together, our results argue against single-axis literature-search evaluation: recall, topical-relevance scoring, ranked-list diversity, and a co-authorship-distance diagnostic each measure complementary properties of citation quality and should be reported jointly.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.