SciFigQual-Bench: 전체 논문 맥락을 고려한 과학적 그림 품질 평가를 위한 벤치마크
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
과학적 이미지는 실험 결과 제시, 시스템 아키텍처 설명, 그리고 비교 주장을 뒷받침하는 데 핵심적인 역할을 합니다. 그러나 기존의 이미지 품질 평가(IQA) 방법은 주로 자연 사진이나 AI 생성 콘텐츠에 맞춰 설계되어 있어, 과학 논문에 직접 적용하기 어렵습니다. 기존 연구들이 학술적 차트를 시각적으로 비교하는 데 그치면서, 캡션과의 일관성, 인용 관련성, 그리고 시각적인 오해 가능성을 검증하지 못했습니다. 이러한 문제를 해결하고자, 우리는 전체 텍스트 맥락을 고려하여 과학적 이미지의 품질을 5가지 측면(명확성, 레이아웃, 캡션 적합성, 문맥 관련성, 오해 위험)에서 평가하는 SciFigQual-Bench를 제안합니다. 이 데이터는 2020년부터 2025년까지 주요 컴퓨터 과학 학회 논문들을 포함하며, 6,308개의 이미지가 여러 분야 전문가에 의해 독립적으로 5가지 측면으로 평가되어, 표준화된 주석으로 통합되었습니다. 기존의 과학적 그림 벤치마크와 달리, 저희 데이터셋은 각 이미지와 연결된 캡션, 인용 문장, 그리고 논문 맥락을 함께 제공합니다. 이 벤치마크에 대한 자동화된 평가를 가능하게 하기 위해, 우리는 여러 모달 증거를 수집하고 통합하여 감사 가능한 정교한 평가를 수행하는 SFQ-Agent라는 단계별 교차 모달 평가 프레임워크를 설계했습니다. GPT-5.6-Sol을 탑재한 SFQ-Agent (F3)는 테스트 데이터셋 eval1200에서 다른 모델보다 낮은 평균 절대 오차(0.418)와 높은 일관성(93.4%)을 보였습니다. 이는 기존의 직접 평가 방법 및 시각 언어 모델 기반의 보조 평가 방식보다 우수한 성능입니다.
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve auditable and refined scoring through the collection and fusion of modal evidence. Multiple mainstream large models were evaluated on the test subset eval1200, and SFQ-Agent (F3) equipped with GPT-5.6-Sol achieved the lowest overall average absolute error (0.418) and the highest consistency rate (93.4%), consistently outperforming both direct evaluation and auxiliary (Sidecar) visual language model evaluation schemes.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.