CrystalXRD-Bench: 다양한 결정성 물질에 대한 XRD 피크 인덱싱을 위한 비전-언어 모델 성능 평가
CrystalXRD-Bench: Benchmarking Vision-Language Models for XRD Peak Indexing Across Diverse Crystalline Materials
분말 X선 회절 패턴으로부터 밀러 지수(Miller index)를 식별하는 것은 기존의 멀티모달 벤치마크에서 검증되지 않은 능력을 요구합니다. 모델은 렌더링된 과학적 곡선에서 미세한 피크 위치를 읽은 다음, 이 정보를 다단계 결정학적 추론과 연결해야 합니다. 본 논문에서는 CrystalXRD-Bench라는 250개의 샘플로 구성된 벤치마크를 소개합니다. 이는 10개의 공개 결정학 데이터베이스를 기반으로 하며, 단일 작업(즉, XRD 패턴에서 가장 강한 피크에 기여하는 전체 HKL 세트 복구)을 수행하도록 설계되었습니다. 각 샘플은 렌더링된 XRD 이미지와 원본 CIF 텍스트 및 화학식을 함께 제공하므로, 시각적 추출 오류와 추론 오류를 비교하여 분석할 수 있습니다. 우리는 7개의 비전-언어 모델을 평가했습니다. 가장 높은 자카드(Jaccard) 점수는 0.5888 (GPT-5.4)이었으며 정확도 일치율은 37.6%였습니다. 그러나 7개 모델 중 6개가 자카드 점수 0.50 미만으로, 이 작업이 아직 해결되지 않았음을 보여줍니다. 오류 패턴은 체계적으로 다양합니다. 두 개의 피크가 있는 경우 특히 취약하며, 회수(recall) 중심의 모델은 HKL을 과도하게 예측하여 커버리지를 높입니다. 또한 CIF 텍스트에 대한 접근성이 결정학적 계산 능력 격차를 해소하지 못하는 것으로 나타났습니다. 본 벤치마크는 모델 순위를 제공할 뿐만 아니라, 현재 비전-언어 모델이 정량적인 과학적 그림에서 실패하는 조건을 식별합니다. 모든 데이터 및 평가 코드는 공개적으로 이용 가능합니다.
Miller-index identification from powder XRD patterns requires capabilities untested by existing multimodal benchmarks: the model must read a narrow peak location from a rendered scientific curve and then connect that observation to multi-step crystallographic reasoning. We introduce CrystalXRD-Bench, a 250-sample benchmark built from 10 public crystallographic databases for a single task: recover the full set of HKLs contributing to the highest-intensity peak in an XRD pattern. Each sample pairs the rendered XRD image with the source CIF text and chemical formula, so visual extraction errors and reasoning errors can be examined side by side. We evaluate seven vision-language models. The best Jaccard score is 0.5888 (GPT-5.4) with an exact-match rate of 37.6%, yet six of seven models remain below Jaccard 0.50; the task is far from solved. Error patterns vary systematically: double-peak cases are especially brittle, recall-heavy models gain coverage by over-predicting HKLs, and access to CIF text does not close the gap in crystallographic calculation. Alongside model rankings, the benchmark identifies the conditions under which current VLMs fail on quantitative scientific figures. All data and evaluation code will be publicly available.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.