2608.03918v1 Aug 04, 2026 cs.CV

언제, 어디를 봐야 할까: 효율적인 장편 비디오 이해를 위한 적응형 시각 증거 스케줄링

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Keqian Li
Keqian Li
Citations: 32,377
h-index: 9
Xuanzhe Liu
Xuanzhe Liu
Citations: 1,488
h-index: 13
Jiayu Chen
Jiayu Chen
Citations: 41
h-index: 4
Zihao Zheng
Zihao Zheng
Citations: 34
h-index: 3
Maoliang Li
Maoliang Li
Citations: 46
h-index: 4
Xiang Chen
Xiang Chen
Citations: 49
h-index: 4
H. Zou
H. Zou
Citations: 11
h-index: 3
Hengyi Zhang
Hengyi Zhang
Citations: 0
h-index: 0

효율적인 장편 비디오 이해는 제한된 수의 프레임을 선택하여 시각적 증거로 사용하는 비전-언어 모델(VLM)이 추론하는 것을 요구합니다. 기존의 관련성 기반 방법은 고정된 프레임 할당량과 후보 풀을 사용한 정적인 일회성 선택에 의존하는 반면, 에이전트 기반 스케줄러는 다단계 추론 및 대화형 검색을 통해 적응성을 달성하지만 계산 비용이 많이 듭니다. 본 논문에서는 학습 과정이 필요 없는 저비용의 쿼리 적응형 시각 증거 스케줄링 프레임워크인 EcoFrame을 제안합니다. EcoFrame은 VLM의 추론 피드백을 활용하여 프레임 할당량을 늘릴 시점과 추가 후보 증거를 검색할 위치를 결정합니다. 구체적으로, 엔트로피 게이트된 예산 스케줄링은 출력 불확실성을 사용하여 현재 증거가 충분하면 조기에 중단하거나 그렇지 않은 경우 점진적으로 프레임 할당량을 확장합니다. 동시에, 어텐션 가이드형 후보 제안은 프레임 수준의 어텐션을 시간적 우선순위로 변환하여 정보가 풍부한 영역에서 집중적인 지역 검색을 가능하게 하면서, 어텐션이 분산될 경우 전체적인 범위를 유지합니다. Video-MME, LongVideoBench 및 MLVU 데이터셋에 대한 실험 결과, EcoFrame은 다양한 VLM 아키텍처에서 더 나은 정확도-효율성 균형을 달성하는 것으로 나타났습니다. Qwen2.5-VL 모델에서 EcoFrame은 평균 정확도가 64.4로, BOLT의 63.5보다 우수하며, AKS 및 BOLT에 비해 최대 $1.85배$ 빠른 속도를 제공합니다. 에이전트 기반 방법인 A.I.R.과 비교했을 때, EcoFrame은 비슷한 수준의 정확도를 유지하면서 최대 $13.5배$ 더 빠른 추론 속도를 제공합니다. 코드 및 관련 자료는 https://github.com/AK-DREAM/EcoFrame 에서 확인할 수 있습니다.

Original Abstract

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based methods rely on static one-shot selection with fixed frame budgets and candidate pools, while agent-based schedulers achieve adaptivity through costly multi-round reasoning and interactive search. We propose EcoFrame, a training-free framework for low-overhead query-adaptive visual evidence scheduling. EcoFrame leverages the VLM's inference feedback to determine when to increase the frame budget and where to search for additional candidate evidence. Specifically, entropy-gated budget scheduling uses output uncertainty to stop early when the current evidence is sufficient or progressively expand the frame budget otherwise. Meanwhile, attention-guided candidate proposal converts frame-level attention into a temporal prior, enabling dense local search in informative regions while preserving global coverage when attention is diffuse. Experiments on Video-MME, LongVideoBench, and MLVU demonstrate that EcoFrame achieves a better accuracy--efficiency trade-off across multiple VLM backbones. On Qwen2.5-VL, EcoFrame achieves an average accuracy of 64.4, surpassing BOLT at 63.5, while providing a $1.85\times$ speedup over AKS and BOLT. Compared with the agent-based A.I.R., EcoFrame maintains comparable accuracy with up to a $13.5\times$ inference speedup. Code will be available at https://github.com/AK-DREAM/EcoFrame.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!