2607.28590v1 Jul 30, 2026 cs.CV

VAD: 다중 모드 온-폴리시 증류에서 시각적 증거를 활용한 대상 재구성

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

Wenxiang Jiao
Wenxiang Jiao
Citations: 96
h-index: 4
Weinan Zhang
Weinan Zhang
Citations: 25
h-index: 3
Weiwen Liu
Weiwen Liu
Citations: 124
h-index: 6
Yong Yu
Yong Yu
Citations: 44
h-index: 3
Jianghao Lin
Jianghao Lin
Shanghai Jiao Tong University
Citations: 1,555
h-index: 20
Qingyao Li
Qingyao Li
Citations: 185
h-index: 6
Shuai Shao
Shuai Shao
Citations: 114
h-index: 3
Zhengxi Lu
Zhengxi Lu
Citations: 290
h-index: 6
Yuan Lu
Yuan Lu
Citations: 44
h-index: 3
Zhiyuan Yao
Zhiyuan Yao
Citations: 34
h-index: 2
Kangning Zhang
Kangning Zhang
Citations: 189
h-index: 5
Yixing Li
Yixing Li
Citations: 45
h-index: 4

다중 모드 온-폴리시 증류(OPD)는 가이드된 교사의 정보를 이용하여 학생이 생성한 경로를 지도함으로써 세분화된 시각적 지식을 전달합니다. 그러나 다음 토큰 수정은 다양한 정보가 혼합되어 있으며, 여기에는 시각 신호뿐만 아니라 언어적 사전 지식 및 교사 모델에 특정한 효과까지 포함됩니다. 핵심 과제는 어떤 수정이 단순히 증류의 강도나 방식을 결정하는 것이 아니라, 실제로 시각적 증거에 의해 뒷받침되는지를 추정하는 것입니다. 본 논문에서는 Visual Attribution Distillation (VAD)이라는 새로운 방법을 제안합니다. VAD는 교사의 수정을 부분적으로 설명할 수 있는 시각적으로 관련된 부분을 추정하는 반사실적 대상 재구성 알고리즘입니다. 각 학생이 생성한 접두사(prefix)에 대해, VAD는 관련 증거가 존재하는 경우와 없는 경우 모두 동일한 교사 모델을 사용하여 평가합니다. 그 결과로 나타나는 중심화된 로그 확률의 변화는 'ut'라는 서명된 값으로 표현되며, 이는 시각적 증거의 방향성을 나타냅니다. 이 값은 어떤 토큰 후보를 지지하거나 반박하는지에 대한 정보를 제공합니다. VAD는 원래 수정 사항을 이 추정된 방향성에 투영하여 개입에 따른 구성 요소와 그 외 부분(잔차)을 분리하고, 이를 바탕으로 학생의 데이터를 기반으로 한 대상(target)을 재구성합니다. 학습 과정에서, 재구성된 대상은 주요 지도 신호 역할을 하며, 가이드된 교사는 약한 정규화 역할을 수행합니다. 40억 및 90억 파라미터 규모의 여섯 가지 세분화된 시각적 벤치마크에서 VAD는 직접적인 가이드 증류 방법과 시각적 이점 가중치를 사용하는 방법보다 더 우수한 성능을 보였습니다. 토큰 수준 분석 및 제어 대상 분석 결과, 추정된 방향성에 따른 구성 요소는 작업 관련 시각적 수정에 더욱 풍부한 정보를 담고 있으며, 특히 오답을 반박하는 경우 더욱 강력한 대상 변화를 유도합니다. 이러한 결과는 반사실적 대상 재구성이 다양한 정보가 혼합된 지도 방식에 대한 효과적인 대안임을 뒷받침합니다.

Original Abstract

Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!