가장 뛰어난 교사가 항상 최선의 교사는 아니다: 학생 중심의 답변 선택
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
최근 LLM 학습은 합성 응답, 추론 과정 및 도구 사용 시연 등 교사(teacher)가 생성한 데이터를 활용하는 경향이 있습니다. 현재 관행에서는 종종 가장 높은 성능을 보이는 교사를 선택하여 학생 학습 데이터를 생성하며, 이는 암묵적으로 교사의 시험 성적을 교육 품질의 지표로 간주합니다. 본 연구에서는 이러한 가정에 대한 반례를 제시합니다. 동일한 질문에 대해 여러 교사가 올바른 답변을 제공하더라도, 가장 뛰어난 교사의 답변이 반드시 특정 학생에게 최적의 학습 자료가 아닐 수 있습니다. 이러한 문제를 해결하기 위해, 본 연구는 학생 중심의 답변 선택(Student-Centric Answer Sampling, SCAS)이라는 프레임워크를 제안합니다. 이 방법은 검증된 교사 생성 답변 중에서 예상되는 학생 중심 학습 비용을 기준으로 답변을 선택합니다. 토큰 단위 기울기 분해를 기반으로, 우리는 이러한 비용에 대한 효율적인 추정치를 도출하고, 이를 활용하여 학습 과정에서 답변 선택을 안내합니다. 30개의 교사 모델, 6개의 학생 기본 모델 및 8가지 작업에 대한 실험 결과, SCAS는 일관되게 학생의 성능을 향상시키는 것으로 나타났습니다. 이는 효과적인 지식 전달이 단순히 교사의 능력뿐만 아니라 현재 학생에게 적합한 학습 자료를 우선적으로 고려해야 함을 시사합니다.
LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations. Current practice often chooses the highest-performing teacher to generate student training data, implicitly treating teacher test performance as a proxy for teaching quality. We show that this assumption can fail: even when multiple teachers provide correct answers to the same question, the answer from the strongest teacher is not necessarily the best supervision for a given student. To address this gap, we propose Student-Centric Answer Sampling (SCAS), a framework that selects from verified teacher-generated answers according to their estimated student-centric learning cost. Motivated by a token-wise gradient decomposition, we derive an efficient forward-only proxy for this cost and use it to guide answer selection during training. Experiments across 30 teacher models, 6 student base models, and 8 tasks show that SCAS consistently improves student performance, suggesting that effective distillation should prioritize supervision matched to the current student rather than teacher strength alone.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.