2607.28154v1 Jul 30, 2026 cs.CV

OPLD: 온 정책 잠재 증류를 통한 다중 모드 추론

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

Tianyang Xu
Tianyang Xu
Citations: 51
h-index: 2
Shoutai Zhu
Shoutai Zhu
Citations: 35
h-index: 3
Bingchuan Sun
Bingchuan Sun
Citations: 0
h-index: 0
Mingyuan Xu
Mingyuan Xu
Citations: 0
h-index: 0
Yu Liu
Yu Liu
Citations: 0
h-index: 0
Qinzhen Guo
Qinzhen Guo
Citations: 0
h-index: 0

다중 모드 체인 오브 소트(Chain-of-Thought, CoT)는 중간 추론 과정에 추가적인 시각 정보를 통합하여 시각적 추론 능력을 향상시킵니다. 그러나 기존 방법은 외부적으로 정의된 추론 경로와 시각 작업에 제약되어 유연하고 추상적인 시각적 사고 능력을 개발하는 데 한계가 있습니다. 최근에는 잠재 변수를 활용한 접근 방식이 중간 계산을 연속적인 표현으로 통합하여 유망한 방향을 제시했습니다. 하지만 기존의 시각-잠재 변수 방법은 주로 압축된 추가 시각 특징과 정렬함으로써 잠재 상태를 지도하며, 이를 시각적 관찰의 대리 값으로 취급합니다. 결과적으로 이러한 방법들은 제공된 증거는 포착하지만 다중 모드 CoT에 의해 유도되는 추상적인 추론 과정을 완전히 내면화하지 못합니다. 본 논문에서는 OPLD(On-Policy Latent Distillation)라는 간단한 프레임워크를 제안합니다. 이 프레임워크는 특권적인 다중 모드 CoT에서 유도된 추론 능력을 잠재 추론 표현으로 전달합니다. 다양한 다중 모드 벤치마크에서의 광범위한 실험 결과, OPLD가 기존의 잠재 추론 방법보다 일관되게 우수한 성능을 보이며, 여러 벤치마크에서 최고 수준의 성능을 달성함을 보여줍니다. 이러한 결과는 잠재 표현을 추론 과정 수준에서 지도하는 것이 기존의 특징 수준 정렬 방식보다 다중 모드 잠재 추론에 더 효과적인 패러다임을 제공한다는 것을 시사합니다.

Original Abstract

Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!