짝 데이터 없이 수행하는 모달 간 지식 증류: 이론적 기반 및 알고리즘
Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm
모달 간 지식 증류(CMKD)는 한 종류의 데이터(예: 이미지)로 학습된 (큰) 교사 모델이 다른 종류의 데이터(예: 텍스트/오디오)를 기반으로 구축되는 (작은) 학생 모델을 어떻게 안내할 수 있는지 연구하는 분야입니다. 기존의 CMKD 방법들은 종종 정렬된 의미를 가진 쌍의 다중 모달 데이터를 필요로 하지만, 이러한 쌍의 데이터를 얻는 것은 비용이 많이 들고 비현실적입니다. 이러한 제한점을 극복하기 위해, 우리는 쌍의 데이터가 없는 더욱 어려운 환경을 위한 새로운 CMKD 프레임워크를 개발했습니다. 특히, 우리는 교사 모델과 학생 모델 간의 모달 간 분포 관계를 확립하여 효과적인 증류를 지배하는 두 가지 기본적인 요소인 특징 정렬 및 레이블 정렬을 밝혀냈습니다. 이러한 요소들은 각각 표현 수준과 예측 분포 수준에서 모달 간의 의미적 불일치를 특성화합니다. 이러한 통찰력을 바탕으로, 우리는 개별 샘플이 아닌 분포를 정렬하여 효과적인 모달 간 지식 증류를 가능하게 하는 이론적 보장을 제공하는 체계적인 프레임워크를 제안했습니다. 광범위한 다중 모달 벤치마크에서의 실험 결과는 우리 프레임워크가 쌍의 데이터와 짝 없는 데이터 환경 모두에서 매우 효과적이며, 기존 연구보다 성능이 크게 향상됨을 보여줍니다.
Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e.g., images) can guide a (smaller) student model building on another type of data (e.g., text/audio). Existing CMKD methods often require paired multi-modal data with aligned semantics, but obtaining such paired data are often costly and impractical. To mitigate this limitation, we develop a new CMKD framework for the more challenging setting where paired data are unavailable. In particular, we establish a cross-modal distributional relationship between teacher and student models, which reveals two fundamental quantities governing effective distillation: feature alignment and label alignment. These quantities characterize semantic discrepancy between modalities at the levels of representation and prediction distributions, respectively. Motivated by this insight, we propose a principled framework, with theoretical guarantees, that enables effective cross-modal knowledge distillation by aligning distributions rather than individual samples. Extensive experiments across a wide range of multimodal benchmarks show that our framework is highly effective in both unpaired and paired data settings, improving significantly over prior work.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.