교사 모델 정보 기반 혼합 사전 지식 증류
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors
지식 증류는 모델 압축을 위한 강력한 방법으로, 대규모 언어 모델을 포함한 복잡한 심층 학습 모델(교사)의 효율적인 배포를 가능하게 합니다. 그러나 이 기술의 근본적인 통계적 메커니즘은 여전히 명확하지 않으며, 특히 다양한 교사 모델의 전문성이 요구되는 실제 환경에서는 불확실성 평가가 종종 간과됩니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 베이지안 프레임워크 내에서 여러 교사 모델로부터 학습하는 지식 증류 방법인 '멀티 티처 베이지안 지식 증류(MT-BKD)'를 제안합니다. 우리의 접근 방식은 베이지안 추론을 활용하여 증류 과정에 내재된 불확실성을 포착합니다. 우리는 교사 모델로부터 얻은 외부 지식과 작업별 훈련 데이터를 통합하는 교사 정보 기반 사전(teacher-informed prior)을 도입함으로써, 더 나은 일반화 성능, 안정성 및 확장성을 제공합니다. 또한, 엔트로피 기반 가중치 조정 메커니즘을 통해 각 교사의 영향력을 적응적으로 조절하여 학생 모델이 다양한 전문 지식을 효과적으로 결합할 수 있도록 합니다. MT-BKD는 학생 모델의 학습 과정을 해석 가능하게 하고 예측 정확도를 향상시키며 불확실성 정량화를 제공합니다. 우리는 MT-BKD를 합성 데이터와 실제 작업 모두에서 검증했으며, 단백질 세포 내 위치 예측 및 이미지 분류 작업을 포함했습니다. 실험 결과는 개선된 성능과 견고한 불확실성 정량화를 보여주며, 우리의 MT-BKD 프레임워크의 강점을 강조합니다.
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.