하나의 순위, 모든 예산: 장편 비디오 이해를 위한 마트료시카 증거-맥락 프레임 선택
One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
프레임 선택은 대규모 다중 모달 모델(LMM)을 장편 비디오에 적용하는 데 필수적입니다. 이는 심각한 프레임 중복과 제한된 컨텍스트 창으로 인해 발생합니다. 적절한 프레임 예산은 하위 작업, 추론 요구 사항 및 지연 시간 제약 조건에 따라 달라지므로, 실용적인 선택기는 다양한 예산을 지원해야 합니다. 그러나 기존 방법은 일반적으로 각 사전 정의된 예산에 대해 최적화된 프레임 부분집합을 사용합니다. 따라서 예산이 변경되면 이전에 선택된 증거가 대체되는 반면 점진적으로 보강되지 않습니다. 고정된 점수로 프레임을 순위화하면 다양한 예산 간에 선행 부분을 재사용할 수 있지만, 이는 서로 다른 순위의 구별된 역할을 무시합니다. 본 논문에서는 장편 비디오 프레임 선택 문제를 마트료시카 순위 문제로 정의합니다. 즉, 작은 접두부는 쿼리에 조건부인 증거를 집중시키고, 점차 더 큰 접두부는 이러한 증거를 유지하면서 광범위한 시간적 맥락을 추가하는 단일 우선순위 시퀀스를 구성합니다. 이러한 순위를 효율적으로 구축하는 것은 자체로 어려운 문제입니다. 왜냐하면 장편 비디오에서 밀집 샘플링을 수행하고 프레임-쿼리 관련성을 평가하는 데 상당한 오버헤드가 발생하기 때문입니다. 따라서 본 논문에서는 Matryoshka Evidence-to-Context (MEC) 프레임 선택이라는 학습이 필요 없는 프레임워크를 제안합니다. 이 프레임워크는 재사용 가능한 희소 비디오 인덱스를 구축하고, 희소 탐색 및 로컬 확대 기능을 통해 후보를 검색하며, 위치에 적응하는 순위를 탐욕적으로 구성합니다. 초기 위치는 증거를 강조하고, 후기 위치는 시각적 다양성을 유지하면서 점진적으로 시간적 범위를 선호합니다. 따라서 단일 순위를 원하는 대상 예산으로 잘라내어 선택기를 다시 실행할 필요가 없습니다. 4개의 벤치마크와 6가지 프레임 예산을 사용하여 MEC은 균일 샘플링보다 평균 정확도를 3.77% 향상시키고, 강력한 최첨단 선택기와 동등한 성능을 보이며, 전체 선택 지연 시간을 47.37-51.19% 단축합니다.
Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However, existing methods typically optimize an isolated frame subset for each predefined budget: when the budget changes, previously selected evidence may be replaced rather than progressively augmented. Ranking frames by a fixed score would allow prefix reuse across budgets, but it ignores the distinct roles of different ranking positions. In this paper, we formulate long-video frame selection as a Matryoshka ranking problem: constructing a single priority sequence whose small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context. Efficiently constructing such a ranking is itself challenging, as densely sampling long videos and evaluating frame-query relevance incurs substantial overhead. We therefore introduce Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that builds a reusable sparse video index, discovers candidates through sparse probing and local zooming, and greedily constructs a position-adaptive ranking: early positions emphasize evidence; later positions progressively favor temporal coverage while preserving visual diversity. A single ranking can thus be truncated to any target budget without rerunning the selector. Across four benchmarks and six frame budgets, MEC improves average accuracy over uniform sampling by 3.77 percentage points, matches strong state-of-the-art selectors, and reduces end-to-end selection latency by 47.37-51.19%.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.