OPD-V: 모달리티 균형을 고려한 시각적 온-폴리시 자기 증류
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
온-폴리시 자기 증류(OPSD)는 다중 모드 대규모 언어 모델(MLLM)의 시각적 추론 능력을 향상시키는 데 널리 사용되는 후처리 기술입니다. 기존 방법들은 다양한 입력 소스에서 얻은 추가 정보를 활용하여 자기 증류 과정을 안내합니다. 그러나 이러한 설계는 MLLM 추론에 내재된 '모달리티 불균형' 문제를 간과합니다. 특히, 텍스트 정보가 생성 과정에서 우세하게 작용할 경우 모델은 다중 모드 입력을 충분히 통합하지 못합니다. 결과적으로, 신중하게 설계된 추가 정보는 제대로 활용되지 못하고, 이는 OPSD의 효과를 제한합니다. 이러한 한계점을 분석하기 위해, 우리는 '확대 이미지'를 사용하여 높은 모달리티 불균형을 보이는 긍정적 교사 모델과 '마스크 이미지'를 사용하여 낮은 모달리티 불균형을 보이는 부정적 교사 모델을 구성했습니다. 이들의 추론 정확도 및 토큰 로짓 값의 변화는 '모달리티 균형' 자체가 중요한 추가 정보가 될 수 있음을 보여줍니다. 이러한 발견에 영감을 받아, 우리는 긍정적/부정적 교사 모델을 통해 모달리티 균형 정보를 활용하는 시각적 OPSD 프레임워크인 OPD-V를 제안합니다. 긍정적인 모달리티 균형 로짓 마진은 '모달리티 균형 신뢰 영역'을 정의하며, 이를 통해 자기 증류에 사용되는 온-폴리시 토큰을 선택합니다. 6개의 벤치마크, 4가지 MLLM 기반 모델, 및 5가지 후처리 방법을 사용하여 실험한 결과, OPD-V는 일관되게 추론 성능을 향상시키면서도 학습 비용을 감소시키는 것으로 나타났습니다.
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalance, a challenge inherent to MLLM reasoning. When textual information dominates generation, the model cannot fully integrate its multimodal input. Consequently, carefully designed privileged information remains underused, limiting the effectiveness of OPSD. To examine this limitation, we construct a Positive Teacher with the Zoom-In Image and a Negative Teacher with the Mask Image, which exhibit different degrees of Modality Imbalance. Changes in their reasoning correctness and token logits reveal that Modality Balance can itself serve as privileged information. Motivated by this finding, we introduce OPD-V, a visual OPSD paradigm that instantiates such information through the Positive Teacher and Negative Teacher. Positive Modality-Balance Logits Margins define a Modality-Balance Trust Region that selects the on-policy tokens used for self-distillation. Experiments across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods show that OPD-V consistently improves reasoning performance while reducing training cost.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.