2606.17507v1 Jun 16, 2026 cs.AI

교육 분야의 LLM 기반 채점 시스템: 교육 과정에 기반한 채점 파이프라인

LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

Xiwei Xu
Xiwei Xu
Citations: 342
h-index: 11
Wenjie Zhang
Wenjie Zhang
Citations: 37
h-index: 4
Chen Wang
Chen Wang
Citations: 5
h-index: 1
Jacky Jiang
Jacky Jiang
Citations: 0
h-index: 0
Phil Yang
Phil Yang
Citations: 0
h-index: 0
Qian Fu
Qian Fu
Citations: 2
h-index: 1
Mohan Dhall
Mohan Dhall
Citations: 7
h-index: 1
Liming Zhu
Liming Zhu
Citations: 28
h-index: 3

생성형 AI와 대규모 언어 모델(LLM)은 질문 생성 및 자동 평가에 점점 더 많이 활용되고 있습니다. 그러나 고위험 시험 준비를 위해 LLM을 사용하는 것은 단순한 프롬프트 엔지니어링 이상의 노력이 필요하며, 교육 기관에서 발행하는 승인된 교육 과정 자료 및 채점 지침과 체계적으로 모델의 출력을 연결하는 소프트웨어 파이프라인이 요구됩니다. 본 논문에서는 대학 입학 시험 준비를 지원하기 위해 산업 파트너와 공동 개발한, 교육 과정에 기반하고 구성 가능한 LLM 기반 채점 파이프라인을 제시합니다. 이 파이프라인은 질문의 관련 주제, 하위 주제 및 인지적 요구 사항을 식별하고, LLM의 판단을 뒷받침할 수 있는 검증 가능하고 승인된 문맥 정보를 제공합니다. 교육 과정의 목표는 지정된 동사 및 결과, 성취 수준 설명, 용어 정의 및 채점 지침 원칙과 같은 구체적인 교육 과정 자료를 통해 구현됩니다. 단계별 LLM 워크플로우를 사용하여 먼저 질문에 특화된 평가 기준을 생성하여 성능에 대한 구조적 기대를 포착하고, 학생 답변에 점수를 할당하는 데 사용되는 채점 기준을 도출하고 평가합니다. 이러한 설계는 일관성, 투명성을 향상시키고 공식적인 채점 방식을 준수하도록 합니다. 예비 평가는 제안된 LLM 기반 채점 파이프라인이 인간 튜터와 비교 가능한 채점 결과를 제공하며, 동시에 승인된 교육 과정 자료 및 채점 기준에 더 추적 가능한 근거를 제시한다는 것을 보여줍니다. 또한 이 파이프라인은 온라인 학습 플랫폼에 통합되었으며, 초기 배포 데이터를 통해 운영 사용 및 수동 수정 사항에 대한 초기 통찰력을 얻을 수 있습니다.

Original Abstract

Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment. However, deploying LLMs in preparation for high-stakes exams requires more than prompt engineering; it demands software pipelines that systematically ground model outputs in authorised curriculum artefacts and marking guidelines issued by education authorities. This paper presents a curriculum-grounded, configurable LLM-as-Judge pipeline for question-level marking, co-developed with an industrial partner, to support exam preparation for university admission. The pipeline identifies the relevant topics, subtopics, and cognitive demand of a question, and assembles verifiable and authorised context to support LLM judgement. Curriculum intent is operationalised through concrete syllabus artefacts, including prescribed verbs and outcomes, performance band descriptors, glossary definitions, and marking-guideline principles. A staged LLM workflow is employed to first generate question-specific rubrics, capturing structured expectations of performance, and then derive and evaluate marking criteria used to allocate marks to student responses. This design improves consistency, transparency, and alignment with official marking practices. Preliminary evaluation shows that the proposed LLM-as-Judge pipeline delivers marking outcomes comparable to human tutors, while yielding justifications that are more traceable to authorised curriculum artefacts and marking standards. The pipeline has also been integrated into an online study platform, where early deployment data provide initial insights into operational usage and manual overrides.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!