2605.26680v1 May 26, 2026 cs.CV

DynFrame: 동적 프레임 증강을 통한 추론 기반 다중 모드 프레임워크 - 복잡한 비디오 이해를 위한 적응형 접근 방식

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding

Jinlong Liu
Jinlong Liu
Citations: 98
h-index: 3
Wanggui He
Wanggui He
Citations: 979
h-index: 14
Pipei Huang
Pipei Huang
Citations: 1,937
h-index: 7
Longxiang Zhang
Longxiang Zhang
Citations: 47
h-index: 2
Yan Xia
Yan Xia
Citations: 13
h-index: 1
Zhenhao Peng
Zhenhao Peng
Citations: 0
h-index: 0
Weilong Dai
Weilong Dai
Citations: 57
h-index: 4
Hao Tang
Hao Tang
Citations: 0
h-index: 0
Le Zhang
Le Zhang
Mila-Quebec AI Institute
Citations: 189
h-index: 9
Mushui Liu
Mushui Liu
Citations: 627
h-index: 13
Peng Zhang
Peng Zhang
Citations: 22
h-index: 3
Guanghao Zhang
Guanghao Zhang
Citations: 201
h-index: 7
Hao Jiang
Hao Jiang
Citations: 94
h-index: 2

최근의 비디오 다중 모드 대규모 언어 모델(MLLM)은 단계별 추론과 필요한 경우 시각적 증거 검색을 결합하여, 모델이 추론 과정에서 관련 비디오 세그먼트를 다시 참조할 수 있도록 합니다. 그러나 기존의 비디오 기반 사고 시스템에는 다음과 같은 두 가지 구조적인 문제가 남아 있습니다. (i) 샘플링 밀도는 학습 가능한 결정 요소가 아닙니다. 기존 방법은 모델이 어디를 봐야 할지 결정하도록 하지만, 프레임 단위의 샘플링 속도는 대부분 고정되어 있습니다. 그 결과, 세밀한 증거는 종종 반복적인 검색을 통해 얻어지므로, 추론 컨텍스트 길이가 증가하고 학습 난이도가 높아집니다. (ii) 검색과 답변 생성은 일반적으로 단일 트래젝토리 수준의 장점을 기준으로 최적화되므로, '어디를 봐야 할지'를 나타내는 토큰과 '어떻게 답변할지'를 나타내는 토큰이 모두 동일한 가중치를 받게 되지만, 실제로 한쪽만 정확할 수도 있습니다. 이러한 문제점을 해결하기 위해, 우리는 시간 윈도우와 샘플링 밀도를 단일 자동 회귀 과정 내에서 기본 토큰으로 출력하는 프레임워크인 DynFrame을 제안합니다. 이 학습 가능한 스팬-밀도 검색은 단일 검색 단계로 다중 수준의 증거를 확보할 수 있도록 합니다. 위에 제시된 토큰화된 검색 인터페이스를 기반으로, 우리는 각 실행(rollout)을 검색 경계에서 분리하고 역할별 토큰 수준의 장점을 부여하는 Segment-Decoupled GRPO (SD-GRPO)를 추가적으로 제안합니다. DynFrame-4B는 선별된 DM-CoT-74k 및 DM-RL-45k 데이터셋으로 학습되었으며, 6개의 벤치마크(NExT-GQA, Charades-STA, ActivityNet-MR, Video-MME, MLVU, LVBench)에서 강력한 7B-8B 모델과 경쟁력을 보입니다. 또한 DynFrame-8B는 대부분의 측정 지표에서 새로운 최고 성능을 달성합니다. 코드 및 관련 자료는 https://github.com/zhangguanghao523/DynFrame 에서 확인할 수 있습니다.

Original Abstract

Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video segments during inference. However, two structural gaps remain in existing thinking-with-video systems. (i) Sampling density is not a learnable decision: existing methods may let the model decide where to look, but the per-window frame rate is largely fixed. As a result, fine-grained evidence is often recovered through repeated retrieval calls, which increases inference context length and training difficulty. (ii) Retrieval and answer generation are usually optimized with a single trajectory-level advantage, so the "where to look" tokens and the "how to answer" tokens receive the same credit even when one is correct and the other is not. To address these gaps, we present DynFrame, a framework that emits the temporal window and the sampling density as native tokens within a single autoregressive pass. This learnable span-density retrieval enables acquiring multi-granularity evidence with a single retrieval step. Based on the above tokenized retrieval interface, we further introduce Segment-Decoupled GRPO (SD-GRPO), which splits each rollout at the retrieval boundary and assigns role-specific token-level advantages, separately crediting the sampling decision and the answer. Trained on the curated DM-CoT-74k and DM-RL-45k, DynFrame-4B is competitive with strong 7B-8B baselines across six benchmarks (NExT-GQA, Charades-STA, ActivityNet-MR, Video-MME, MLVU, LVBench), and DynFrame-8B sets new state-of-the-art on most metrics. Code is available at https://github.com/zhangguanghao523/DynFrame.

0 Citations
0 Influential
30.4657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!