교육 과정 맞춤화: 동적 데이터-모델 호환성을 통한 학생 중심 추론 증류
Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility
추론 증류는 대규모 언어 모델(LLM)의 복잡한 추론 능력을 소형 모델로 이전하는 기술이지만, 이 성공 여부는 학습 데이터가 학생 모델과 얼마나 잘 일치하는지에 달려 있습니다. 본 논문에서는 데이터-모델 호환성(DMC) 지표를 소개합니다. DMC는 학생 모델에 대한 추론 증류에 적합한 데이터 세트를 평가하는 데 사용될 수 있습니다. DMC는 데이터 품질, 상대적 난이도 및 학생 모델의 역량을 종합적으로 고려하여 평가를 제공합니다. 우리는 다음 두 가지 관점에서 DMC의 효과성을 검증했습니다: (1) DMC는 추론 증류 성능과 높은 상관관계를 보입니다; 그리고 (2) DMC를 데이터 선택 기준점으로 사용하면 추론 증류 성능이 향상됩니다. 이러한 결과는 여러 학생 모델 및 작업에 걸쳐 일관되게 나타났습니다. 또한, 각 데이터 세트의 DMC 값이 학습 중에 동적으로 변하기 때문에, DMC를 기준으로 데이터 세트를 동적으로 선택하면 성능을 더욱 향상시킬 수 있다는 것을 실험을 통해 확인했습니다.
Reasoning distillation transfers complex reasoning abilities from large language models (LLMs) to smaller ones, yet its success depends on how well the training data align with the student model. This paper introduces the Data-Model Compatibility (DMC) metric, which can be used to assess the suitability of a dataset for reasoning distillation on a student model. DMC provides an assessment by jointly considering data quality, relative difficulty, and student capability. We validated the effectiveness of DMC from two perspectives: (1) DMC exhibits a strong correlation with reasoning distillation performance; and (2) using DMC as the criterion for data selection leads to improved reasoning distillation performance. Both findings are consistently demonstrated across multiple student models and tasks. Moreover, since the DMC of each dataset dynamically changes during training, our experiments demonstrate that dynamically selecting datasets based on DMC can further enhance performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.