2607.28098v1 Jul 29, 2026 cs.AI

SciDataSailor: 심층 과학 데이터 탐색

SciDataSailor: Deep Scientific Data Exploring

Chun-dong Song
Chun-dong Song
Citations: 31
h-index: 2
Runkai Zhao
Runkai Zhao
Citations: 25
h-index: 1
J. Rao
J. Rao
Citations: 1
h-index: 1
Yicheng Qiu
Yicheng Qiu
Citations: 0
h-index: 0
Chi Zhang
Chi Zhang
Citations: 1
h-index: 1

과학 데이터 세트는 일반적으로 계층 구조로 구성된 저장소이며, 이 저장소에는 서로 다른 형식과 의존성을 가진 다양한 파일들이 포함되어 있어, 이러한 데이터를 검토하고 통합하며 분석하는 데 많은 노력이 필요하고 전문 지식이 요구됩니다. 대규모 언어 모델(LLM) 에이전트가 계획, 추론 및 도구 사용 능력에서 상당한 발전을 이루었지만, 기존 연구에서는 실행 가능한 환경을 통해 실제 과학 데이터 자산과 상호 작용하는 에이전트의 능력이 간과되었습니다. 본 논문에서는 Deep Scientific Data Exploration이라는 에이전트 기반 작업 패러다임을 소개합니다. 이 패러다임에서 에이전트는 저장소를 탐색하고, 다양한 파일 및 스키마를 해석하며, 분석을 실행하고, 여러 파일을 연결하여 얻은 정보를 통합하고, 실행된 관찰 결과를 바탕으로 결론을 도출합니다. 이러한 패러다임을 구현하기 위해, 본 논문에서는 SciDataSailor라는 프레임워크를 제시합니다. 이 프레임워크는 광범위한 탐색과 목표 지향적인 활용 사이의 균형을 맞추어 도구와 상호 작용하는 경로를 합성합니다. SciDataSailor는 몬테카를로 트리 검색(MCTS)을 사용하여 경로 합성을 구현하며, 네 가지 작업별 메커니즘을 포함합니다: 난이도에 따른 탐색 초기화 방법, 이중 피드백 기반의 우선순위 부여, 계층적 전략-to-tool 액션 생성, 그리고 엔트로피 기반 분기. 본 프레임워크를 사용하여 SciDataSailor-SFT-2K는 지도 학습을 위한 모델로, SciDataSailor-Bench는 평가를 위한 벤치마크로 구축되었습니다. SciDataSailor-Bench는 생명 과학, 지구 과학 및 물리 과학 분야의 27개 데이터 세트에 걸쳐 627개의 메타 정보 요약 작업과 586개의 과학 질문 답변 작업을 포함합니다.

Original Abstract

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments. We introduce Deep Scientific Data Exploration, an agentic task paradigm in which agents navigate repositories, interpret heterogeneous files and schemas, execute analyses, integrate cross-file evidence, and produce conclusions grounded in executed observations. To operationalize this paradigm, we present SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation. SciDataSailor instantiates trajectory synthesis as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms: difficulty-stratified exploration seeds, dual-feedback first-play urgency, hierarchical strategy-to-tool action generation, and entropy-guided branching. Using this framework, we construct SciDataSailor-SFT-2K for supervised fine-tuning and SciDataSailor-Bench for evaluation, with the latter comprising 627 meta-information summarization tasks and 586 scientific question-answering tasks across 27 datasets spanning the life, earth, and physical sciences.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!