2603.29139v1 Mar 31, 2026 cs.AI

SciVisAgentBench: 과학 데이터 분석 및 시각화 에이전트 평가를 위한 벤치마크

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

Nathaniel Gorski
Nathaniel Gorski
Citations: 22
h-index: 3
Shusen Liu
Shusen Liu
Citations: 53
h-index: 5
Bei Wang
Bei Wang
Citations: 17
h-index: 3
Kuangshi Ai
Kuangshi Ai
Citations: 53
h-index: 5
Haichao Miao
Haichao Miao
Citations: 129
h-index: 7
Kaiyuan Tang
Kaiyuan Tang
Citations: 125
h-index: 7
Jianxin Sun
Jianxin Sun
Citations: 392
h-index: 6
Guoxi Liu
Guoxi Liu
Citations: 30
h-index: 4
H. Ingólfsson
H. Ingólfsson
Citations: 8,069
h-index: 37
David Lenz
David Lenz
Citations: 53
h-index: 5
Hongfeng Yu
Hongfeng Yu
Citations: 84
h-index: 6
Teja Leburu
Teja Leburu
Citations: 5
h-index: 1
Michael Molash
Michael Molash
Citations: 5
h-index: 1
T. Peterka
T. Peterka
Citations: 34
h-index: 4
Chaoli Wang
Chaoli Wang
Citations: 1,308
h-index: 20
Hanqi Guo
Hanqi Guo
Citations: 54
h-index: 5

최근 대규모 언어 모델(LLM)의 발전으로 자연어 의도를 실행 가능한 과학 시각화(SciVis) 작업으로 변환하는 에이전트 시스템이 등장했습니다. 그러나 이러한 새로운 SciVis 에이전트를 실제 다단계 분석 환경에서 평가할 수 있는 체계적이고 재현 가능한 벤치마크는 아직 부족합니다. 본 논문에서는 과학 데이터 분석 및 시각화 에이전트를 평가하기 위한 포괄적이고 확장 가능한 벤치마크인 SciVisAgentBench를 제시합니다. 벤치마크는 응용 분야, 데이터 유형, 복잡성 수준, 시각화 작업의 네 가지 차원을 포괄하는 체계적인 분류 체계를 기반으로 하며, 현재는 다양한 SciVis 시나리오를 다루는 108개의 전문가가 설계한 사례로 구성되어 있습니다. 신뢰성 있는 평가를 위해, 우리는 LLM 기반 판단과 이미지 기반 지표, 코드 검사기, 규칙 기반 검증기, 사례별 평가기를 포함하는 다중 모드 중심 평가 파이프라인을 도입했습니다. 또한, 12명의 SciVis 전문가를 대상으로 인간 평가자와 LLM 평가자 간의 일관성을 검증하는 타당성 연구를 수행했습니다. 이 프레임워크를 사용하여 대표적인 SciVis 에이전트와 범용 코딩 에이전트를 평가하여 초기 기준선을 설정하고 성능 격차를 파악했습니다. SciVisAgentBench는 체계적인 비교를 지원하고, 실패 요인을 진단하며, 에이전트 기반 SciVis 분야의 발전을 촉진하기 위한 지속적인 벤치마크로 설계되었습니다. 벤치마크는 https://scivisagentbench.github.io/ 에서 이용 가능합니다.

Original Abstract

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We also conduct a validity study with 12 SciVis experts to examine the agreement between human and LLM judges. Using this framework, we evaluate representative SciVis agents and general-purpose coding agents to establish initial baselines and reveal capability gaps. SciVisAgentBench is designed as a living benchmark to support systematic comparison, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is available at https://scivisagentbench.github.io/.

8 Citations
0 Influential
18.5 Altmetric
100.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!