2605.29833v1 May 28, 2026 cs.AI

OmniMatBench: 인간 전문가의 검증을 거친 다중 모드 추론 벤치마크 - 19개 재료 과학 분야

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

Qian Tan
Qian Tan
Citations: 214
h-index: 3
Lei Bai
Lei Bai
Citations: 75
h-index: 4
Weida Wang
Weida Wang
Shanghai AI Laboratory
Citations: 189
h-index: 9
Zhuo Yang
Zhuo Yang
Citations: 13
h-index: 2
Jiaqing Xie
Jiaqing Xie
Citations: 54
h-index: 4
Tianfan Fu
Tianfan Fu
Citations: 7
h-index: 1
Yuqiang Li
Yuqiang Li
Citations: 29
h-index: 4
Lu Chen
Lu Chen
Citations: 460
h-index: 12
Wanli Ouyang
Wanli Ouyang
Citations: 4,084
h-index: 20
Wanhao Liu
Wanhao Liu
Citations: 119
h-index: 5
Ran Sun
Ran Sun
Citations: 297
h-index: 4
Jue Wang
Jue Wang
Citations: 160
h-index: 2
Xin Chen
Xin Chen
Citations: 178
h-index: 7

다중 모드 언어 모델이 과학 연구에서 점점 더 중요한 역할을 수행함에 따라, 재료 과학은 그 융합적 특성, 다중 모드 데이터, 그리고 응용 중심적인 성격으로 인해 중요한 테스트 환경을 제공합니다. 그러나 기존의 재료 관련 벤치마크는 주로 물성 예측, 지식 질의응답, 또는 특징 이해에 초점을 맞추고 있으며, 재료 지식을 활용한 보다 광범위한 추론 과정은 상대적으로 간과되어 왔습니다. 이러한 격차를 해소하기 위해, 우리는 인간 전문가의 검증을 거친 다중 모드 추론 벤치마크인 OmniMatBench를 제안합니다. OmniMatBench는 19개의 재료 과학 하위 분야에 걸쳐 3,171개의 전문가가 선별한 질의응답 및 계산 문제로 구성되어 있으며, 여기에는 기초 재료 지식, 구조 및 공학 재료, 재료 가공 및 제조, 그리고 기능성 및 응용 재료가 포함됩니다. 우리는 13개의 오픈 소스 및 상업용 다중 모드 언어 모델을 평가한 결과, 가장 성능이 좋은 모델조차도 0.372의 낮은 전반적인 점수를 기록했으며, 이는 현재 재료 과학 분야의 추론 능력에 상당한 격차가 존재함을 보여줍니다. 추가 분석 결과, 하위 분야별로 큰 차이가 있으며, 고정된 추론 방법, 불균형한 재료 지식, 그리고 수식, 검색, 코드 지원 환경에서 제한적인 고급 지식 활용이 관찰되었습니다. OmniMatBench는 현재 다중 모드 언어 모델의 능력과 한계에 대한 중요한 통찰력을 제공하며, 재료 과학 연구 분야에서 신뢰할 수 있는 인공지능 어시스턴트 개발을 위한 기반을 마련합니다.

Original Abstract

As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdisciplinary, multimodal, and application-driven nature. However, existing materials benchmarks mainly focus on property prediction, knowledge QA, or characterization understanding, leaving the broader reasoning process from materials knowledge to application underexplored. To fill this gap, we present OmniMatBench, a human-calibrated multimodal reasoning benchmark for materials science. OmniMatBench contains 3,171 expert-curated QA and calculation problems across 19 materials-science subfields, spanning fundamental materials knowledge, structural and engineering materials, materials processing and manufacturing, and functional and applied materials. We evaluate 13 open-source and closed-source MLLMs and find that the best model achieves only a 0.372 overall score, revealing a substantial gap in current materials-science reasoning. Further analysis shows strong variation across subfields, fixed reasoning heuristics, uneven materials knowledge, and limited high-level knowledge application under formula-, retrieval-, and code-assisted settings. OmniMatBench provides crucial insights into the capabilities and limitations of current MLLMs and establishes a foundation for reliable AI assistants in materials-science research.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!