LLM 기반 튜터 모델 평가 시 '도움이 되는' 정도를 교육적 신호로 재해석하는 연구: 사전 등록된 감사 결과 분석
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
LLM 기반 튜터링은 측정상의 문제를 야기합니다. 일반적인 '도움을 주는' 기준이 직접적인 답변 제공과 교육적 지침을 구별할 수 있는가? 본 연구는 사전 등록된 방식으로 이 신호의 유효성을 검증합니다. 세 가지 튜터 모델을 기준으로, 동일한 기반 모델을 사용하고 하나의 고정된 수준 낮은 시뮬레이션 학생과 연결된 대화형 정책과 교육적 정책을 비교합니다. 결정적인 감지기는 답변 노출 여부와 다음 발언에 대한 학생의 독립성 정도를 측정합니다. Claude Opus 4.8은 조건에 영향을 받지 않는 주요 평가 모델로 사용되었습니다. Opus의 점수가 확정된 후, GPT-5.6 Sol은 동일한 1,179개의 확인적인 답변 단계 튜터 대화 내용을 대상으로, 사후적으로 '도움을 주는' 정도와 교육적 기준을 사용하여 신뢰성 검증을 수행했습니다. 주요 모델에서 Opus를 기준으로 평가했을 때, 정책 간의 '도움이 되는' 정도 차이는 통계적으로 유의미하지 않지만, 교육적 측면에서는 뚜렷한 순위 차이를 보였습니다(Cliff's $|δ| = 0.10$ vs. $1.0$). 두 평가 모델을 비교했을 때, 교육적 기준에 따른 대비는 일관성을 유지하는 반면, '도움이 되는' 정도의 순서는 평가 모델에 따라 달라졌습니다(세 가지 모델 중 두 곳에서 역전). Opus만을 사용한 분석 결과, 주요 모델 하위 정책들은 평균적으로 평가된 교육적 측면에서 2.3점 차이를 보이지만, 평균적으로 평가된 '도움이 되는' 정도는 0.25점 이내의 범위에 있었습니다. 별도로, 답변을 노출하는 대화 이후에는 모든 모델에서 학생들의 독립적인 활동이 줄어드는 경향을 보였으며, 이는 평가 모델과 관계없이 동일하게 나타났습니다. 이러한 통제된 환경에서, 일반적인 '도움이 되는' 정도는 신뢰할 수 있는 교육적 신호가 아닙니다. 튜터 평가 시에는 교육적 측면에 특화된 기준과 함께 결정적인 프로세스 측정 지표를 함께 사용해야 합니다.
LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|δ|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.