2605.26781v1 May 26, 2026 cs.AI

LiveK12Bench: 거대 멀티모달 모델이 정말로 고등학교 수준의 시험을 정복했는가?

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

Gang Liu
Gang Liu
Citations: 43
h-index: 4
Xiaohan Wang
Xiaohan Wang
Citations: 101
h-index: 6
Ming Yin
Ming Yin
Citations: 445
h-index: 2
Yilin Zhao
Yilin Zhao
Citations: 23
h-index: 3
Dian Li
Dian Li
Citations: 110
h-index: 3

최첨단 거대 멀티모달 모델(LMM)은 K-12 교육 과정의 추론 과제에서 뛰어난 성능을 보여주며, 지능형 튜터로서 큰 잠재력을 가지고 있습니다. 이러한 잠재력을 실현하기 위해서는 모델이 실제 시험 환경을 효과적으로 처리할 수 있어야 하지만, 현재 대부분의 벤치마크는 실제 시험 환경의 복잡성을 제대로 반영하지 못합니다. 특히, 대부분의 데이터셋은 정적이며, 데이터 오염에 취약하고, 종종 제한된 모달리티, 학문 분야 및 평가 기준으로 구성됩니다. 이러한 문제점을 해결하기 위해, 우리는 현실적인 시험 시나리오에서 LMM의 추론 능력을 평가하도록 설계된 동적이고 종합적인 멀티학문 벤치마크인 LiveK12Bench를 소개합니다. LiveK12Bench는 수학, 물리, 화학, 생물학 분야에 걸쳐 2,000개 이상의 검증된 문제로 구성되어 있으며, 최신 실제 시험 문제를 기반으로 구축되었으며 시간이 지남에 따라 확장될 예정입니다. 우리의 프레임워크는 다음과 같은 핵심 혁신을 포함합니다: 1) 데이터 유출을 방지하기 위해 최신 시험지를 지속적으로 수집하고 분석하는 자동화된 파이프라인; 그리고 2) 정확하고 효율적인 추론 경로를 사용하여 자율적으로 전체 과정을 완료할 수 있는 능력을 평가하는 새로운 '모의 시험' 평가 시스템. 12개의 LMM에 대한 광범위한 실험 결과, 최첨단 모델조차도 실제 시험과 유사한 제약 조건 하에서 상당한 성능 저하를 경험한다는 것을 보여줍니다. 예를 들어, GPT-5의 점수는 프로세스의 엄격성과 효율성을 동시에 평가할 때 79점에서 53점으로 감소했습니다. 우리의 연구 결과는 복잡한 시각적 레이아웃에 대한 민감성과 같이 중요한 취약점을 드러내며, 이상적인 추론 능력과 실제 교육 적용 가능성 간의 격차를 강조합니다. 코드와 데이터셋은 공개적으로 제공됩니다.

Original Abstract

Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors. Realizing this potential requires models to navigate real-world examinations effectively, yet most existing benchmarks fail to capture the complexity of authentic testing environments. Specifically, most datasets are static, prone to data contamination, and are often confined to restricted modalities, disciplines, and evaluation criteria. To address these issues, we introduce LiveK12Bench, a dynamic, holistic, multi-disciplinary benchmark designed to evaluate the reasoning abilities of LMMs in realistic examination scenarios. LiveK12Bench comprises 2K+ verified questions spanning Mathematics, Physics, Chemistry, and Biology, sourced from the latest real-world exam papers and designed to grow over time. Our framework features several core innovations: 1) featuring an automated pipeline that continuously ingests and parses the latest examination papers to mitigate data leakage; and 2) proposing a novel `Mock Exam' evaluation scheme, which assesses the ability to complete end-to-end exams autonomously with accurate and efficient reasoning paths. Extensive experiments on 12 LMMs reveal that advanced models suffer substantial performance degradation under exam-realistic constraints: GPT-5's score drops from 79 to 53 (out of 100) when process rigor and efficiency are jointly evaluated. Our findings expose critical vulnerabilities, such as sensitivity to complex visual layouts, highlighting the gap between idealized reasoning capabilities and true educational readiness. Both code and dataset are publicly available.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!