2607.20918v1 Jul 23, 2026 cs.AI

OPOD: 온-폴리시 다중 모드 지식 증류

OPOD: On-Policy Omni Distillation

Yuyang Hu
Yuyang Hu
GSAI
Citations: 203
h-index: 3
Zhicheng Dou
Zhicheng Dou
Citations: 3,020
h-index: 28
Tong Zhao
Tong Zhao
Citations: 31
h-index: 4
Haibo Shi
Haibo Shi
Citations: 201
h-index: 6
Yutao Zhu
Yutao Zhu
Citations: 140
h-index: 4
Reed Li
Reed Li
Citations: 0
h-index: 0
Yu Lu
Yu Lu
Citations: 87
h-index: 5

다중 모드 모델은 텍스트, 이미지 및 오디오를 하나의 시스템에서 처리할 수 있지만, 이러한 다양한 능력을 동시에 향상시키는 것은 여전히 어렵습니다. 풀링된 다중 모드 데이터로 단일 모델을 학습하는 경우, 종종 개별 모드에 특화된 모델만큼의 성능을 내지 못합니다. 온-폴리시 지식 증류(OPD)는 이러한 전문화된 모델들을 결합하는 방법으로, 학생 모델이 응답을 생성하면, 교사 모델은 동일한 응답을 평가하여 학생 모델이 실제로 생성하는 행동으로부터 직접 학습하도록 합니다. 그러나 여러 개의 교사를 사용하면 상반되는 가이드라인이 도입되어 특정 모드의 성능 향상이 다른 모드의 성능 저하를 초래할 수 있습니다. 본 논문에서는 각 학생 모델의 응답을 해당 텍스트, 이미지 또는 오디오 교사에게 연결하는 온-폴리시 다중 모드 지식 증류(OPOD) 방법을 제시합니다. OPOD는 교사의 가이드라인을 학생 모델이 할당한 확률보다 높은 확률을 가진 토큰에만 적용하고, 각 모드 교사의 영향력을 독립적으로 조정하며, 연결된 교사에게 최종 답변뿐만 아니라 추론 과정이 정답을 뒷받침하는지 여부를 평가하도록 합니다. 12개의 벤치마크와 3가지 크기의 모델에서 OPOD는 모든 규모에서 가장 높은 평균 점수를 달성했으며, 각각 70.8점, 51.7점 및 46.2점을 기록하여 가장 강력한 비교 모델보다 2.1점, 1.8점 및 1.7점 높았습니다. 30B 모델의 경우, OPOD는 풀링된 다중 모드 데이터로 공동 학습을 수행한 기본 모델과 유사 모델 모두를 능가했으며, 개별 전문화 모델이 포함되더라도 11개의 벤치마크에서 1위 또는 2위를 차지했습니다. 학습 후에는 개별 전문화 모델은 제거되어 하나의 배포 가능한 다중 모드 모델만 남게 됩니다. 이러한 결과는 모드별 교사를 조정하는 것이 공유 모델의 성능을 향상시키면서도 모드 간 균형을 유지하는 효과적인 방법임을 보여줍니다.

Original Abstract

Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult. Training a single model on pooled multimodal data often fails to match models specialized for individual modalities. On-policy distillation (OPD) offers a way to combine such specialists: the student generates a response, and a teacher evaluates that same response, so the student learns directly from behaviors it actually produces. Yet using several teachers can introduce competing guidance and improve one modality at the expense of another. We present On-Policy Omni Distillation (OPOD), which routes each student response to the matching text, image, or audio teacher. OPOD keeps teacher guidance only on tokens where the teacher assigns a higher probability than the student, adjusts the influence of each modality teacher independently during training, and asks the routed teacher to assess both the final answer and whether the reasoning supports the correct answer. Across twelve benchmarks and three backbone sizes, OPOD achieves the best average score at every scale, reaching 70.8, 51.7, and 46.2 and exceeding the strongest comparator by 2.1, 1.8, and 1.7 points. On the 30B model, it outperforms both the base model and a counterpart post-trained jointly on pooled multimodal data on all twelve benchmarks, and ranks first or second on eleven even when the individual specialists are included. The specialists are discarded after training, leaving one deployable omni-modal model. These results show that coordinating modality-specific teachers is an effective way to improve a shared model while maintaining cross-modal balance.

0 Citations
0 Influential
14 Altmetric
70.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!