RP-OPSD: 해상도 우위 기반 온라인 자기 증류를 통한 다중 모드 대규모 언어 모델
RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
온라인 자기 증류(OPSD)는 교사 모델만이 가진 특권 정보를 활용하여 학생이 생성한 경로에 대한 밀집적인 토큰 수준의 감독 신호를 제공합니다. 그러나 기존 방법은 종종 검증된 해결 과정, 외부 모델에서 생성된 설명 또는 수동으로 식별된 시각적 증거에 의존하는데, 이는 다중 모드 대규모 언어 모델에 대한 확장 가능한 적용을 제한합니다. 이러한 문제를 해결하기 위해, 저희는 동일 이미지의 고해상도 및 저해상도 뷰 간의 정보 격차를 활용하여 RP-OPSD(Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models)를 제안합니다. 학습 과정에서, 학생 모델은 원래 해상도의 1/4 크기의 이미지로부터 온라인 경로를 생성하고, 교사 모델은 원래 해상도 이미지를 사용하여 감독 신호를 제공합니다. 학생 경로는 고해상도 입력에 대한 교사 모델의 예측 행동을 학습하여 학생 모델이 저해상도 환경에서의 성능을 향상시키고, 이를 원래 해상도의 추론으로 이전하도록 합니다. RP-OPSD는 추가적인 인간 어노테이션이나 해결 과정을 생성하기 위한 외부 모델을 필요로 하지 않으며, 이미지-질문 쌍만 사용합니다. Qwen3.5-9B 모델에 대한 실험 결과, RP-OPSD는 원래 해상도에서 평균 성능이 5.45% 향상되고 OPSD보다 학습 속도가 $1.78 imes$ 빨라졌습니다. 이러한 결과는 해상도 차이가 간단하고 확장 가능한 특권 정보의 원천이 될 수 있으며, 다중 모드 대규모 언어 모델을 위한 온라인 자기 증류에 효과적이고 효율적인 접근 방식을 제공한다는 것을 보여줍니다.
On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solution traces, explanations generated by external models, or manually localized visual evidence, which limits their scalable application to multimodal large language models. To address this issue, we exploit the information gap between high- and low-resolution views of the same image and propose RP-OPSD (Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models). During training, the student policy generates on-policy trajectories from images at one-quarter of the original resolution, while the teacher policy provides supervision using the original-resolution images. By minimizing the divergence between their output distributions along the student trajectories, the student learns the predictive behavior of the teacher under high-resolution inputs, thereby strengthening its low-resolution capability and transferring the learned improvement to original-resolution inference. RP-OPSD requires neither additional human annotations nor external models to generate solution traces but only image--question pairs. Experiments on Qwen3.5-9B show that RP-OPSD achieves a 5.45\% relative improvement in average performance at the original resolution and a $1.78\times$ training speedup over OPSD. These results demonstrate that resolution differences can serve as a simple and scalable source of privileged information, providing an effective and efficient approach to on-policy self-distillation for multimodal large language models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.