2608.05780v1 Aug 06, 2026 cs.CV

효율적인 장편 비디오 이해를 위한 증거 기반 동적 시각 선택 방법

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

Feng Chen
Feng Chen
Citations: 84
h-index: 6
Bo Zhang
Bo Zhang
Citations: 838
h-index: 3
Changsheng Li
Changsheng Li
Citations: 38
h-index: 3
Yinjie Lei
Yinjie Lei
Citations: 29
h-index: 3
Wenxin Wang
Wenxin Wang
Citations: 0
h-index: 0
Zhihao Zhang
Zhihao Zhang
Citations: 0
h-index: 0
Zixuan Wang
Zixuan Wang
Citations: 29
h-index: 3

최근 MLLM(Large Multimodal Language Model) 기반의 장편 비디오 이해 기술 발전은 관련 프레임을 선택함으로써 추론 시간 동안 발생하는 계산 비용과 제한된 컨텍스트 길이를 줄이는 데 기여했습니다. 그러나 기존 방식들은 주로 외부 프록시 스코어링 및 경직된 휴리스틱 규칙에 의존하기 때문에, 대상 MLLM의 고유한 증거와 일치하지 못하고 비균일한 시공간 정보 밀도를 수용하는 데 어려움을 겪습니다. 본 논문에서는 대상 MLLM의 내부 어텐션 증거를 기반으로 하는 정교한 동적 시각 선택 프레임워크인 EviSelect를 제안합니다. 저희 방법은 구조화된 사전 지식을 활용하여 시각적 증거를 효율적으로 탐색하고, 분포에 대한 인지적인 동적 샘플링을 안내합니다. 구체적으로, 저희는 대상 MLLM의 어텐션 맵을 고도로 압축된 시각 입력과 희소 어텐션을 사용하여 효율적으로 근사하며, 이는 전체 버전에 잘 맞도록 설계되었습니다. 이 사전 지식에서 파생된 세 가지 상호 보완적인 어텐션 구성 요소를 기반으로, 저희는 가벼운 선택기를 설계했습니다. 이 선택기는 쿼리 관련 타임스탬프를 정확하게 식별할 뿐만 아니라, 로컬 샘플링 비율과 공간 해상도를 적응적으로 조정합니다. 증거에 조건화된 시공간 샘플링을 가능하게 하기 위해, 저희는 선택기를 확률적 정책으로 정의하고 GRPO(Group Relative Policy Optimization)를 통해 최적화합니다. 그룹 상대 비교를 통해 낮은 시각 비용에서 정확한 예측을 보상함으로써, 저희 방법은 정책이 각 비디오의 정보 밀도에 따라 계산 자원을 동적으로 할당하도록 장려합니다. 세 가지 장편 비디오 이해 벤치마크에서 EviSelect는 기존 방법보다 우수한 성능을 달성했으며, 선택된 시각적 토큰 수를 약 50% 줄이고 전체 속도를 3.9배 향상시켰습니다.

Original Abstract

Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM's intrinsic evidence and failing to accommodate the non-uniform spatiotemporal information density. In this paper, we propose a fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence. Our method efficiently probes visual evidence via sparse prefilling as a structured prior to guide distribution-aware dynamic sampling. Specifically, we efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart. Conditioned on three complementary attention components derived from this prior, we design a lightweight selector that not only precisely locates query-relevant timestamps but also adaptively adjusts the local sampling rate and spatial resolution. To enable evidence-conditioned spatiotemporal sampling, we formulate the selector as a stochastic policy and optimize it via GRPO under a joint accuracy--efficiency reward. By rewarding correct predictions under lower visual cost through group-relative comparisons, our method encourages the policy to allocate computation dynamically according to the information density of each video. Across three long video understanding benchmarks, EviSelect achieves superior performance compared to existing methods while reducing selected visual tokens by about 50\% and achieving a 3.9x end-to-end speedup.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!