2606.13148v1 Jun 11, 2026 cs.AI

TerraBench: 에이전트가 다양한 지구 시스템 데이터를 이해할 수 있는가?

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

Numan Saeed
Numan Saeed
Citations: 367
h-index: 11
F. Maani
F. Maani
Citations: 183
h-index: 8
Muhammad Umer Sheikh
Muhammad Umer Sheikh
Citations: 40
h-index: 3
Salman Khan
Salman Khan
Citations: 83
h-index: 2
Dat Nguyen
Dat Nguyen
Citations: 54
h-index: 2
Thao Nguyen
Thao Nguyen
Citations: 59
h-index: 1
Muhammad Haris Khan
Muhammad Haris Khan
Citations: 33
h-index: 2
Huy M. Le
Huy M. Le
Citations: 5
h-index: 1

기후 및 환경 의사 결정은 점점 더 다양한 입력 데이터, 즉 격자 형태의 물리 데이터, 위성 이미지, 지리 공간 정보 및 시뮬레이션 결과 등을 포괄적으로 분석해야 하는 상황으로 발전하고 있습니다. 날씨 및 기후 예측 모델은 뛰어난 예측 성능을 보이지만 자연어 인터랙션을 제공하지 못하며, 대규모 언어 모델(LLM)은 자연어로 추론할 수 있지만 고차원 지구 시스템 데이터를 직접 처리할 수 없습니다. 결과적으로, 지구 과학 분야의 실제 연구 워크플로우는 충분히 지원받지 못하고 있습니다. 본 논문에서는 TerraAgent라는 ReAct 스타일의 실행 가능한 프레임워크를 기반으로 구축된, 지구 과학 추론을 위한 벤치마크인 TerraBench를 소개합니다. TerraAgent는 LLM의 계획 기능을 환경 데이터 검색, 지리 공간 처리, 시뮬레이션 및 증거 기반 계산을 위한 과학적 도구와 결합하여 추론, 도구 호출 및 관찰을 통합합니다. TerraBench는 지구 관측 이미지 분석, 격자 데이터, GIS 추론 및 시뮬레이션을 하나의 실행 가능한 인터페이스로 통합하는 반면, 기존 벤치마크는 이러한 기능을 개별적인 제한된 작업으로 분리했습니다. 또한, 본 연구는 최초로 프로세스 수준의 도구 사용 지표와 허용 오차를 고려한 수치 점수를 결합하여 평가합니다. TerraBench는 세 가지 트랙(기본, 시뮬레이션 기반, 문서 기반 검증)과 8개의 응용 분야에 걸쳐 총 403개의 광범위한 에이전트 관련 작업으로 구성되어 있으며, 24,500개의 검증된 실행 단계를 포함합니다. 이러한 결과는 신뢰할 수 있는 지구 과학 에이전트가 다양한 워크플로우를 조정하고, 도구를 정확하게 매개변수화하며, 데이터의 출처 정보를 보존하는 기능을 갖추어야 함을 시사합니다.

Original Abstract

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!