2608.03206v1 Aug 04, 2026 cs.CY

EduClaw-Bench: 시뮬레이션된 학습자를 활용한 교육용 LLM 에이전트를 위한 장기 평가 벤치마크

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

Unggi Lee
Unggi Lee
Citations: 484
h-index: 11
Sookbun Lee
Sookbun Lee
Citations: 40
h-index: 3
Yeil Jeong
Yeil Jeong
Citations: 85
h-index: 5
Eunjoo Lee
Eunjoo Lee
Citations: 1
h-index: 1
Minchul Shin
Minchul Shin
Citations: 0
h-index: 0
H. Kwon
H. Kwon
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 튜터링부터 에세이 채점까지 다양한 교육 분야에서 활용되지만, 각각은 단일 작업에 대한 개별적인 솔루션이며, 최근 들어 이러한 개별 솔루션들이 학습 관리 시스템(LMS) 내에서 작동하는 에이전트로 통합되기 시작했습니다. 그러나 튜터링은 장기적인 과정으로, 학습자의 실력 향상은 하루나 몇 번의 상호작용으로는 이루어지지 않고 며칠 또는 몇 주에 걸쳐 진행됩니다. 따라서, 지속적인 관계 속에서 에이전트 튜터를 평가하는 벤치마크는 아직 존재하지 않습니다. 본 연구에서는 EduClaw-Bench를 소개합니다. 이는 에이전트 튜터가 지식 추적(KT) 기술을 기반으로 한 시뮬레이션된 학습자와 30일 동안 지속적인 관계를 맺도록 설계된 벤치마크입니다. 시뮬레이션된 학습자의 답변은 실제 학생 데이터를 기반으로 학습한 KT 모델에서 파악되는 지식-개념 숙달도를 통해 결정되며, 55개의 다양한 시나리오에서 학습 효과를 측정합니다. 각 에이전트는 세 가지 주요 평가 기준(학습 효과, 응답성 및 유용성)과 두 가지 교육 과정 설계 기준(Gagné와 Rosenshine)에 따라 평가됩니다. 유용성과 교육 과정 설계 기준은 LLM 3개로 구성된 패널 심사위원단이 판단합니다. 세 개의 기본 모델 계층에서 10개의 에이전트 어댑터를 평가한 결과, 단일 계층 및 단회원 평가로는 얻을 수 없는 두 가지 중요한 사실을 발견했습니다. 첫째, 튜터링 품질은 개별적인 요소(기본 모델 또는 에이전트 자체)보다는 이들이 결합된 결과에 더 크게 영향을 받습니다. 둘째, 거의 모든 조합에서 전체 기간 동안 일관된 수준의 양질의 튜터링을 제공하지 못했습니다. 교정 검증($ ext{ECE}=0.049$) 및 실제 수업 환경에서의 필드 연구를 통해 시뮬레이션된 학습자와 측정값이 현실과 유사함을 확인했습니다. 본 연구는 미래 교육을 위한 신뢰할 수 있는 AI 튜터 개발에 기여하는 중요한 단계입니다.

Original Abstract

Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!