2607.20926v1 Jul 23, 2026 cs.AI

SciExplore: 과학 탐색부터 정보 통합까지 자율 에이전트의 평가

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

Weiming Zhang
Weiming Zhang
Citations: 114
h-index: 5
Wenran Liu
Wenran Liu
Citations: 315
h-index: 5
Yanan Sun
Yanan Sun
Citations: 118
h-index: 4
Yinhao Tang
Yinhao Tang
Citations: 92
h-index: 3
Youqing Fang
Youqing Fang
Citations: 1
h-index: 1
Bin Liu
Bin Liu
Citations: 0
h-index: 0
Kuikun Liu
Kuikun Liu
Citations: 965
h-index: 9
Wenwei Zhang
Wenwei Zhang
Citations: 3,155
h-index: 21
Kai Chen
Kai Chen
Citations: 0
h-index: 0

과학 연구는 다양한 출처에 걸쳐 복잡한 정보 검색 및 추론 워크플로우를 포함합니다. 그러나 기존 벤치마크는 주로 일반적인 영역의 정보 검색 또는 정적 과학 질문 답변을 강조하며, 따라서 실제 과학 연구 워크플로우에서 요구되는 핵심 기능을 평가하는 데 실패합니다. 본 논문에서는 LLM과 에이전트의 과학 정보 검색 및 추론 능력을 평가하기 위한 벤치마크인 SciExplore를 소개합니다. SciExplore는 과학 데이터베이스 탐색, 모호한 문헌 검색, 누락된 참고문헌 완성, 여러 출처의 구조화된 지식 종합 등 103개의 전문가가 선별한 작업으로 구성되어 있으며, 이는 개체 수준 추론 및 문서 수준 식별부터 증거 수준 기반 분석 및 도메인 수준 종합에 이르기까지 점진적으로 높은 수준의 능력을 평가합니다. 우리는 SciExplore를 사용하여 최첨단 LLM과 자율 에이전트 10개 이상을 평가한 결과, 작업 복잡성이 증가함에 따라 성능 저하가 심화되고 가장 어려운 구조화된 종합 작업에서는 매우 낮은 정확도를 보이는 상당한 성능 격차가 있음을 확인했습니다. 이러한 결과는 현재 모델 및 에이전트가 실제 과학 정보 검색 시나리오에서 가지고 있는 중요한 한계점을 강조합니다.

Original Abstract

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!