2608.04124v1 Aug 04, 2026 cs.CV

추론에 앞선 인식: 비디오 이해 및 질의응답을 위한 동적 잠재 추론

Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

Junbo Zou
Junbo Zou
Citations: 41
h-index: 4
Haotian Xia
Haotian Xia
Citations: 96
h-index: 6
Zilin Xiao
Zilin Xiao
Rice University
Citations: 63
h-index: 4
Hanjie Chen
Hanjie Chen
Rice University
Citations: 2,036
h-index: 16
Vicente Ordonez
Vicente Ordonez
Citations: 13
h-index: 2

비디오 질의 응답 시스템은 모델이 언어적 질문을 시각적 증거에 연결하고, 필요한 경우 시간 경과에 따라 해당 증거를 기반으로 추론할 수 있도록 해야 합니다. 기존 방법들은 종종 긴 텍스트 기반의 사고 과정을 사용하지만, 많은 질문들은 관련 객체, 동작 또는 프레임을 찾자마자 답변될 수 있습니다. 본 논문에서는 Dynamic Latent Reasoning (DyLaR)을 제안합니다. DyLaR은 먼저 질문을 짧은 인식 잠재 변수 블록(질문에 관련된 시각적 증거를 인코딩하는 연속적인 숨겨진 상태)에 연결한 다음, 답변하기 전에 추론 잠재 변수(잠재 공간에서 이 증거에 대해 추론하는 연속적인 사고)를 추가할지 여부를 적응적으로 결정합니다. DyLaR은 검증된 시각적 증거에 인식 잠재 변수를 연결하고, 검증된 설명 과정을 추론 잠재 변수로 추출함으로써 이러한 동작을 학습하며, 강화 학습을 통해 언제 추론해야 하는지를 더욱 개선합니다. 9개의 비디오 벤치마크와 4가지 멀티모달 언어 모델 백본에서 DyLaR은 동일한 백본을 사용하는 기존 방법보다 평균 정확도를 향상시키면서, 질문당 20개 미만의 토큰을 사용하여 답변을 생성합니다. 예를 들어, Qwen3-VL-4B를 사용할 때, DyLaR은 Qwen3-VL-4B-Thinking 모델의 평균 정확도를 54.0에서 58.2로 향상시키면서, 질문당 응답 길이를 1,220.7 토큰에서 18.5 토큰으로 줄였습니다. 추가적인 실험 결과는 연결된 인식 잠재 변수, 설명 과정 기반의 추론 잠재 변수, 그리고 적응적 라우팅이 각각 정확도를 향상시킨다는 것을 보여줍니다.

Original Abstract

Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the relevant object, action, or frame is localized. We propose Dynamic Latent Reasoning (DyLaR), which first grounds a question in a short block of perception latents (continuous hidden states that encode query-relevant visual evidence), and then adaptively decides whether to append reasoning latents (continuous thoughts that reason over this evidence in latent space) before answering. DyLaR learns this behavior by grounding perception latents in verified visual evidence and distilling verified rationales into reasoning latents, followed by reinforcement learning that further refines when to reason. Across nine video benchmarks and four multimodal language model backbones, DyLaR improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query. On Qwen3-VL-4B, for example, DyLaR improves average accuracy over Qwen3-VL-4B-Thinking from 54.0 to 58.2 while reducing response length from 1,220.7 to 18.5 tokens per query. Ablations further show that grounded perception latents, rationale-supervised reasoning latents, and adaptive routing each improve accuracy.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!