2606.05718v1 Jun 04, 2026 cs.CV

ViCuR: 시각적 단서를 활용한 다중 모드 온폴리시 증류를 위한 회복 가능한 특권

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

Siyuan Liu
Siyuan Liu
Citations: 25
h-index: 3
Kanghui Tian
Kanghui Tian
Citations: 16
h-index: 2
Ziang Yan
Ziang Yan
Citations: 647
h-index: 9
Sheng Xia
Sheng Xia
Citations: 13
h-index: 1
Shuai Dong
Shuai Dong
Citations: 36
h-index: 3
Yi Wang
Yi Wang
Citations: 4,822
h-index: 18

온폴리시 증류(OPD)는 교사의 감독 하에 학생 모델이 자신의 정책에서 샘플링된 경로들을 사용하여 추론 능력을 향상시키는 방법입니다. 다중 모드 추론에서, 흔히 사용되는 방법은 학습 시에만 존재하는 신호(예: 정답 또는 설명)를 관찰하는 특권적인 교사를 사용하는 것입니다. 그러나 이러한 답변 측면의 특권은 학습-테스트 불일치를 야기합니다. 즉, 교사의 감독은 학생에게는 사용할 수 없는 신호에 의존할 수 있으며, 이는 시각적으로 기반한 추론보다는 단순한 모방을 장려할 수 있습니다. 본 논문에서는 ViCuR이라는 시각적으로 기반한 특권적인 교사 증류 프레임워크를 제안합니다. ViCuR은 답변 측면의 특권 대신 입력에 포함된 쿼리와 관련된 시각적 단서(증거)를 활용합니다. 이러한 단서는 추론 시에도 사용 가능한 동일한 시각적 입력을 기반으로 생성되므로, 학생 모델이 해당 증거를 회복할 수 있습니다. 이를 지원하기 위해 ViCuR은 전처리 과정에서 전용 싱크-토큰 크로스-어텐션을 사용하여 관련 시각적 증거를 내부 표현으로 통합하는 경량화된 단서 회복 모듈을 도입합니다. 이 방법은 추론 인터페이스를 변경하지 않으며, 추가적인 단서 생성 손실을 요구하지 않습니다. Qwen3-VL-2B 및 8B 모델을 사용한 7개의 벤치마크에서 ViCuR은 답변 기반 온폴리시 자기 증류 방식보다 전반적인 평균 성능에서 +1.19 및 +1.24의 향상을 보였습니다. 또한, ViCuR은 더 강력한 교사를 사용하는 OPD 방식으로 확장될 수 있으며, 기존 OPD 방식보다 +0.64 및 +1.08의 성능 향상을 보여주며, 8B 규모에서는 일관된 일반화 성능을 나타냅니다. 이러한 결과는 다중 모드 온폴리시 증류에서 교사의 특권 설계가 교사 모델의 강도만큼 중요하다는 것을 보여줍니다.

Original Abstract

On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a common extension is to use a privileged teacher that observes training-time-only signals such as reference answers or rationales. However, such answer-side privilege creates a train-test mismatch: the teacher's supervision may depend on signals unavailable to the student, encouraging shortcut imitation rather than visually grounded reasoning. We propose ViCuR, a visually grounded privileged-teacher distillation framework that replaces answer-side privilege with visual cues (query-related evidence in the input). Because these cues are derived from the same visual input available at inference, their evidence is recoverable by the student. To support this, ViCuR introduces a lightweight cue recovery module that uses dedicated sink-token cross-attention during prefill to aggregate task-relevant visual evidence into an internal representation, without changing the inference interface or requiring auxiliary cue-generation losses. Across seven benchmarks with Qwen3-VL-2B and 8B students, ViCuR consistently improves over answer-based on-policy self-distillation by +1.19 and +1.24 on overall average performance. It also extends naturally to stronger-teacher OPD, surpassing OPD baselines by +0.64 and +1.08, with consistent out-of-domain gains at the 8B scale. These results show that, in multimodal on-policy distillation, the design of teacher privilege is as important as teacher strength.

2 Citations
0 Influential
9 Altmetric
47.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!