과학 시각화 이해도를 위한 다중 모드 대규모 언어 모델 성능 비교 연구
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
다중 모드 대규모 언어 모델(MLLM)은 시각화 자료 해석에 점점 더 많이 활용되고 있지만, 현재의 평가 방법은 주로 차트 중심적이며 과학 시각화(SciVis)에 대한 이해도를 제대로 측정하지 못하는 한계가 있습니다. 본 연구에서는 6개의 MLLM을 대상으로 표준화된 SciVis 이해도 평가 도구를 사용하여 성능을 비교 분석했습니다. 이 평가는 18개의 과학 시각화 자료 및 그림을 기반으로 구성된 49개 문항으로, 8가지 기법과 11가지 유형의 과제를 포함합니다. 폐쇄형 환경에서 3개의 상용 모델과 3개의 오픈 소스 모델을 평가하고, 485명의 인간 참가자 데이터를 사용하여 성능을 비교했습니다. 결과에 따르면 현재 MLLM은 SciVis 이해도 측면에서 일관된 성능을 보이지 않습니다. Gemini는 전체적으로 가장 뛰어난 성능을 보였으며, 일부 평가 항목에서는 인간 평균치를 능가하는 결과를 얻었습니다. 반면 오픈 소스 모델은 전반적으로 인간의 기준 수준 미만으로 평가되었습니다. 모델의 성능은 기법 및 과제 유형에 따라 큰 차이를 보였습니다. 과학적 그림 이해, 검색, 공간 지각 능력 측면에서는 좋은 성과를 보였지만, 텍스처 기반 시각화, 통합형 시각화, 그리고 정량적 추정 능력 측면에서는 어려움을 겪는 것으로 나타났습니다. 오차 분석 결과, 세밀한 정량적 추정, 흐름 방향 해석, 그리고 맥락에 따른 정보 해독에서 반복적인 오류가 발생하는 것을 확인했습니다. 이러한 연구 결과를 바탕으로 SciVis 이해도를 다중 모드 AI 시스템 평가의 필수적인 기준으로 제시합니다. 본 연구의 코드 및 모델 출력 결과는 다음 GitHub 저장소에서 공개적으로 이용할 수 있습니다: https://github.com/patdmp/mllm-scivis-lit-benchmark.
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.