최종 답변을 넘어: 투명한 다중 모드 추론 평가를 위한 CRYSTAL 벤치마크
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
본 논문에서는 6,372개의 데이터 인스턴스로 구성된 진단 벤치마크인 CRYSTAL(Clear Reasoning via Yielded Steps, Traceability, and Logic)을 소개합니다. CRYSTAL은 검증 가능한 중간 단계를 통해 다중 모드 추론을 평가합니다. 우리는 두 가지 상호 보완적인 지표를 제안합니다. 첫째, 의미적 유사성 매칭을 통해 단계별 정밀도와 재현율을 평가하는 Match F1이고, 둘째, 비정형적인 추론 과정을 더욱 감점하는 Ordered Match F1입니다. 참조 데이터는 델파이 방식에 영감을 받은 파이프라인을 통해 구축되었으며, 4개의 독립적인 다중 모드 언어 모델(MLLM)이 추론 경로를 생성하고, 생성된 경로는 의미적 클러스터링을 통해 통합되며, 인간 검토 과정을 거쳐 품질을 검증합니다. 벤치마크 구축 과정에 사용되지 않은 상용 최첨단 시스템을 포함하여 20개의 MLLM을 평가한 결과, 답변 정확도만으로는 파악할 수 없는 체계적인 오류가 발견되었습니다. 이러한 오류에는 광범위한 편향 선택(정밀도가 재현율보다 훨씬 높은 경우), 비단조적인 확장성 문제, 그리고 경쟁 모델 중 어느 것도 일관된 단계 순서를 60% 이상 유지하지 못하는 비정형적인 추론 과정 등이 있습니다. 평가 외에도, 우리는 답변의 정확성과 단계별 일관성을 결합하는 곱셈 보상인 Causal Process Reward (CPR)와, CPR-Curriculum을 제안합니다. CPR-Curriculum은 훈련 과정에서 점진적으로 추론의 난이도를 높이며, 기존의 가법 보상 전략이 실패하는 GRPO(Generalized Reinforcement Proximal Policy Optimization)에서 Match F1을 32% 향상시켜, 수동 단계 주석 없이 추론 능력을 향상시킵니다.
We introduce CRYSTAL (Clear Reasoning via Yielded Steps, Traceability, and Logic), a diagnostic benchmark with 6,372 instances that evaluates multimodal reasoning through verifiable intermediate steps. We propose two complementary metrics: Match F1, which scores step-level precision and recall via semantic similarity matching, and Ordered Match F1, which further penalizes disordered reasoning chains. References are constructed through a Delphi-inspired pipeline in which four independent MLLMs generate trajectories, which are then aggregated via semantic clustering and validated through human quality gates. Evaluation of 20 MLLMs, including commercial frontier systems not used during benchmark construction, reveals systematic failures that are invisible to answer accuracy: universal cherry-picking (precision far exceeds recall), non-monotonic scaling trade-offs, and disordered reasoning in which no competitive model preserves more than 60% of matched steps in the correct order. Beyond evaluation, we propose the Causal Process Reward (CPR), a multiplicative reward that couples answer correctness with step-level alignment, and CPR-Curriculum, which progressively increases reasoning difficulty during training. CPR-Curriculum achieves a 32% improvement in Match F1 via GRPO where additive reward strategies fail, improving reasoning without manual step annotation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.