AI 기반 인간 튜터 평가: 교육 성과와 실제 적용 연계
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
다양한 튜터 교육 플랫폼이 존재하지만, 실제 수행 능력을 기반으로 AI 기반의 튜터 교육 및 평가를 제공하는 곳은 많지 않습니다. 본 연구에서는 실제적인 교육 상황에서의 튜터 역량 향상을 위한 AI 기반 시스템을 제시합니다. 기존 온라인 교육 또는 시뮬레이션을 통해 학습 성과만을 평가하는 플랫폼과는 달리, 본 시스템은 생성형 AI (Gemini-2.5-pro)를 활용하여 실제 튜터링 과정의 기록 내용을 분석하고, 튜터의 기술이 실제 적용으로 어떻게 이어지는지 측정합니다. 수학 과목을 원격으로 가르치는 튜터 86명을 대상으로 시나리오 기반 교육 과정을 진행한 결과, 평균적으로 7.4%의 유의미한 학습 성과 향상이 있었습니다. 혼합 효과 모델 분석 결과, 교육 과정에서의 성과가 실제 튜터링 기록 점수에 유의미하게 영향을 미치는 것으로 나타났으며, 효과 크기는 0.25 SD였습니다. 모델 비교 (AIC/BIC) 결과, 교육 과정에서 개방형 질문 및 객관식 문제에 대한 답변을 종합적으로 고려하는 것이 실제 튜터 성과를 가장 잘 예측했으며, 특히 개방형 질문이 더 높은 예측력을 보였습니다. 추가 분석 결과, 교육 후 튜터들은 자신의 기술을 적용할 수 있는 교육 기회를 경험할 가능성이 유의미하게 증가했습니다 (61.1%에서 68.9%) 또한, 이러한 기회 내에서 실행 품질도 향상되었습니다 (65.5%에서 68.1%). 시계열 분석 결과, 튜터의 개선은 즉각적인 교육 효과보다는 시간이 지남에 따라 점진적으로 나타나는 경향이었습니다. 본 연구는 AI 기반 방법을 통해 튜터 교육과 실제 평가를 연결하는 방안을 제시하며, 투명성과 재현성을 확보하기 위해 관련 데이터셋, AI 프롬프트 및 채점 기준을 공개합니다.
There exist numerous tutor training platforms. However, few provide AI-driven training and evaluation for human tutors based on real-life performance. We present an AI-driven system that assesses both open responses during training and authentic real-life tutoring. Unlike platforms that only assess learning through online training or simulations, our system utilizes Generative AI (Gemini-2.5-pro) to analyze transcriptions of authentic tutoring, measuring the transfer of tutor skills to real-life application. Human tutors instructing students remotely in math (N=86) completed six scenario-based lessons, averaging a significant 7.4% learning gain. Using mixed-effects models across 405 session-to-lesson pairs, we found that training performance significantly predicted real-life transcript scores with an effect size of 0.25 SD. Model comparison (AIC/BIC) indicated averaging open response and multiple choice performance during training predicted real-life tutor performance best, although open responses were comparatively more predictive. Exploratory analysis showed that after training, tutors were significantly more likely to encounter pedagogical opportunities to apply their skills (61.1% to 68.9%) and demonstrated higher execution quality within those opportunities (65.5% to 68.1%). Interrupted time series analysis suggested that these tutor improvements were part of a gradual trend over time rather than an immediate intervention effect of training. We illustrate an AI-driven method to link tutor training with real-life assessment. In doing so, we contribute open datasets, AI prompts, and scoring rubrics to support transparency and reproducibility.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.