2607.28336v1 Jul 30, 2026 cs.AI

보이지 않는 오류 수정: 다중 모드 추론 시스템에서 인지 증류를 위한 신용 할당

Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reasoners

Hongyu Lin
Hongyu Lin
Citations: 4,172
h-index: 29
Feng Xiong
Feng Xiong
Citations: 121
h-index: 5
Leyan Xue
Leyan Xue
Citations: 1
h-index: 1

온라인 학습 기반 증류는 다중 모드 추론 시스템에 대한 밀집적인 지도 정보를 제공하지만, 궤적 수준의 보상은 오답이 인지 단계에서 발생했는지 아니면 후속 추론 단계에서 발생했는지를 판단할 수 없습니다. 여러 추론 과정에서 공유되는 인지 결과를 기반으로 추정되는 인지 성공률(PSR)은 낮은 성공률이 인지 능력 부족과 추론 난이도 중 어느 것에 해당하는지 명확하게 구분하기 어렵기 때문에 모호합니다. 본 연구에서는 다운스트림 오류 및 교사-학생 간 불일치를 상호 보완적인 증거로 활용하여 수정 가능한 인지 실패를 식별하는 라벨 없는 방법인 **인지 수정 증류(PCD)**를 제안합니다. 이 두 가지 증거의 곱은 소프트 AND 게이트를 형성하며, 두 가지 증거 모두 존재할 때만 증류 과정을 강화합니다. 우리는 베이지안 증거 결합을 통해 이러한 규칙을 정당화하고, 어떤 증거라도 없을 때 0이 되는 유일한 정규화된 양선형 게이트가 곱셈이라는 것을 보여줍니다. PCD는 독립적인 인지-추론 시퀀스를 사용하며, 평균 보존 가중치를 사용하여 추론 목표를 변경하지 않습니다. 8개의 벤치마크에서 PCD는 OPD(On-Policy Distillation)를 사용할 때의 8B 2B 모델의 평균 성능을 44.50에서 47.28로 향상시키고, 32B 8B 모델의 성능을 56.94에서 61.22로 향상시켰습니다. 일치된 실험에서 PCD와 독립적인 시퀀스를 제거하면 보정되지 않은 평균 성능이 각각 2.22점과 0.88점 감소했습니다. 따라서 효과적인 다중 모드 증류는 교사가 예측하는 것뿐만 아니라, 인지가 수정의 적절한 대상인지 여부를 식별하는 데 달려 있습니다.

Original Abstract

On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), estimated from multiple reasonings sharing one perception, remains ambiguous because low success conflates perceptual insufficiency with reasoning difficulty. We introduce \textbf{Perception-Correction Distillation (PCD)}, a label-free method that identifies correctable perception failures using downstream failure and teacher--student disagreement as complementary witnesses. Their product, , forms a soft AND gate that strengthens distillation only when both witnesses are present. We motivate this rule through Bayesian evidence combination and show that multiplication is the unique normalized bilinear gate that vanishes when either witness is absent. PCD uses separated perception--reasoning rollouts and mean-preserving weights, leaving the reasoning objective unchanged. Across eight benchmarks, PCD improves the 8B 2B macro average from 44.50 with OPD to 47.28 and the 32B 8B result from 56.94 to 61.22. In matched 2B ablations, removing PCD and separated rollout reduces held-out average by 2.22 and 0.88 points, respectively. Effective multimodal distillation therefore depends not only on what the teacher predicts, but also on identifying when perception is the appropriate target of correction.

0 Citations
0 Influential
14.5 Altmetric
72.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!