HG-Bench: 자동화된 숙제 평가를 위한 다페이지 필기 답안 영역 지칭 벤치마크
HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment
자동화된 숙제 평가는 학생의 답변을 인식하는 것뿐만 아니라, 노이즈가 많고 여러 페이지로 구성된 필기 숙제에서 각 답변과 중간 추론 단계가 어디에 나타나는지를 정확하게 파악하는 데 달려 있습니다. 본 논문에서는 페이지 정보를 고려한 2단계 답안 영역 지칭 평가 환경의 부재를 해결하고자 합니다. 제시된 모델은 숙제 페이지 이미지 시퀀스를 입력받아 완전한 답변 영역과 그 하위 단계를 순서대로 위치시키는 기능을 수행해야 합니다. 우리는 HG-Bench라는 벤치마크를 소개합니다. 이 벤치마크는 1,489,278장의 이미지 풀에서 선별된 500개의 인간이 직접 주석을 달은 초중등학교 숙제 샘플로 구성되어 있으며, 질문 수준과 단계 수준의 경계 상자가 계층적 포함 관계를 통해 연결되어 있습니다. HG-Bench는 페이지 정보를 고려한 평가 프로토콜과 함께 제공되며, 완전한 답변 영역의 위치 정확도(FA)와 단계별 분해 능력(FSm)을 개별적으로 측정합니다. 이를 통해 모델이 단순히 보이는 텍스트를 분석하는 것이 아니라, 학생의 추론 과정의 공간적 구조를 실제로 이해하는지를 파악할 수 있습니다. 최첨단 비공개 API 및 경쟁력 있는 오픈 웨이트 시각-언어 모델(VLM)을 대상으로 한 실험 결과, 단일 모델이 FA에서 55.22% 이상 또는 FSm에서 48.22% 이상의 성능을 보이지 않았습니다. 반면, 약 10,000개의 해당 분야 데이터로 미세 조정된 GLM-4.6V 9B 참조 모델은 각각 74.97/72.26의 성능을 달성했습니다. 이러한 결과는 단계별 필기 영역 지칭이 구체적인 기술 격차를 나타냄을 보여주며, 자동화된 숙제 평가 연구에 대한 재현 가능한 벤치마크, 평가 프로토콜 및 학습된 참조 모델을 제공합니다.
Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate reasoning step appears in noisy, multi-page handwritten work. This paper addresses the missing evaluation setting of page-aware, two-level answer-region grounding: given a sequence of homework page images, a model must localize complete answer regions and their ordered step-level subregions. We introduce HG-Bench, a benchmark of 500 human-annotated K-12 homework samples curated from a 1,489,278-image source pool, with question-level and step-level boxes linked by a hierarchical containment constraint. HG-Bench is paired with a page-aware evaluation protocol that separately measures complete-answer localization (FA) and step-level decomposition (FSm), revealing whether models truly ground the spatial structure of student reasoning rather than merely parse visible text. Across frontier closed-source APIs and competitive open-weight VLMs, no zero-shot system exceeds 55.22% on FA or 48.22% on FSm, while a GLM-4.6V 9B reference model fine-tuned on ~10k in-domain examples reaches 74.97/72.26. These results identify step-level handwritten grounding as a concrete capability gap and provide a reproducible benchmark, evaluation protocol, and trained reference point for future work on automated homework assessment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.