2608.04554v1 Aug 05, 2026 cs.CL

문제 난이도 예측을 위한 시각적 증거 표현: 시각적 텍스트화 및 이미지 기반 모델링

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tianyi Zhou
Tianyi Zhou
Citations: 946
h-index: 11
Han Chen
Han Chen
Citations: 63
h-index: 4
Hong Jiao
Hong Jiao
Citations: 34
h-index: 4

문제의 내용을 분석하여 난이도를 예측하는 것은 충분한 학생 응답 데이터가 없을 때 새롭게 개발된 문제에 대한 초기 추정치를 제공할 수 있습니다. 기존 방법은 일반적으로 문제 제시문과 선택지를 텍스트로 표현합니다. 수학 문제에 시각적 요소가 포함될 경우, 일반적인 방식은 해당 증거를 먼저 텍스트로 변환한 다음 텍스트 예측 모델을 적용하는 것입니다. 본 연구에서는 질문 난이도 예측을 위해 시각적 증거는 어떻게 표현되어야 하는지에 대한 질문을 던집니다. 우리는 문제 텍스트만 사용하는 방법, 시각적 증거를 언어로 표현하는 시각적 텍스트화 방법, 그리고 원본 이미지를 그대로 사용하는 이미지 기반 모델링 방법을 비교합니다. 학생 응답을 통해 난이도가 조정된 Eedi 문제를 사용하여 대규모 언어 모델(LLM)과 시각-언어 모델(VLM)을 직접 학습시켜 난이도 예측을 수행했습니다. 두 가지 시각적 인터페이스 모두 가장 낮은 예측 오차를 보였지만, 최고 성능 시스템의 순위를 명확하게 결정하기는 어려웠습니다. Open-VLM의 텍스트화 방식은 모든 평가된 LLM에서 더 낮은 RMSE(Root Mean Squared Error) 점수를 제공했으며, 더욱 광범위한 적용은 모든 이미지 기반 VLM에서 그러한 결과를 보였습니다. 테스트 단계에서의 추가적인 분석 결과, 예측 성능은 전체 문제 이미지에 의존하는 경향을 보였지만, 추가적인 시각적 요소의 영향력을 명확하게 분리하기는 어려웠습니다. 또한, 두 가지 시각적 인터페이스는 부분적으로 상호 보완적인 오류를 발생시키며 계산 과정에서도 상당한 차이를 보입니다. 따라서 텍스트화 방식만을 실용적인 방법으로 간주해서는 안 되며, 이미지 기반 모델링은 VLM의 적용 방식에 따라 효과적인 대안이 될 수 있습니다.

Original Abstract

Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textualizes that evidence and then applies a text predictor. We ask: how should visual evidence be represented for item difficulty prediction? We compare question text alone, visual textualization, which expresses visual evidence in language, and image-native modeling, which retains the original image. Using Eedi items with difficulty calibrated from student responses, we train large language models (LLMs) and vision-language models (VLMs) directly for difficulty regression. Both visual interfaces achieve the lowest point estimates, although the leading systems cannot be reliably ordered. Open-VLM textualization yields lower RMSE point estimates for all evaluated LLMs, while broader adaptation does so for all image-native VLMs. Test-time interventions show dependence on the paired full-item image, but do not isolate the additional visual component. The two visual interfaces also make partially complementary item-level errors and differ substantially in computational workflow. Thus, textualization should not be treated as the only practical interface: image-native modeling is a competitive alternative whose effectiveness depends on how the VLM is adapted.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!