2608.01660v1 Aug 03, 2026 cs.CV

기반 다지기, 범위 확장 및 정제: 증거 중심 프레임 선택을 통한 장편 비디오 질의응답

Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Runmin Dong
Runmin Dong
Citations: 1,338
h-index: 19

장편 비디오 질의응답은 수천 개의 프레임을 포함하는 비디오에서 제한된 시각-토큰 예산 내에서 중요하면서도 희소한 증거를 식별해야 하는 작업을 의미합니다. 기존 방법들은 대부분 단일 단계로 질의에 관련된 프레임을 선택하거나, 타임스탬프가 있는 텍스트만을 검색 지침으로 활용하여 두 가지 주요 한계를 보입니다. 첫째, 선택된 프레임은 종종 특정 영역의 관련성이 높은 부분에 집중되는 경향이 있으며, 예산이 소진되면 누락된 증거를 복구할 수 없습니다. 둘째, 텍스트와 시각적 증거 간의 연관성이 약합니다. 본 연구에서는 학습 과정이 필요 없는 GCR (Ground, Cover, and Refine) 프레임워크를 제안합니다. 이 프레임워크는 고정된 예산 내에서 프레임을 선택하는 문제를 공동 증거 큐레이션 문제로 정의합니다. '기반 다지기(Ground)' 단계에서는 타임스탬프가 있는 텍스트를 시간적 이벤트로 변환하고, 질의와 관련된 실제 프레임을 선택하며, 각 이벤트를 해당 시간적으로 정렬된 프레임에 매핑합니다. '범위 확장(Cover)' 단계에서는 추가적인 시각적 증거를 위해 직접적인 시각적 기준으로 보완하고, 다양한 맥락을 유지하기 위해 전역 최대 주변 관련성을 적용합니다. '정제(Refine)' 단계에서는 누락된 시간 영역을 재검토하고, 가장 약한 부분을 대체할 수 있는 실제 프레임의 중심값을 사용하되, 중심값이 더 높은 증거 가치를 제공하는 경우에만 수행합니다. GCR은 고정된 개수의 시간순으로 정렬된 프레임을 유지하며, VLM 학습이나 구조적 변경이 필요하지 않습니다. LongVideoBench 및 Video-MME 데이터셋에서 세 가지 7B 모델을 사용하여 프레임 예산을 8, 32, 그리고 64로 설정했을 때, GCR은 장편 비디오 질의응답 성능에서 일관된 개선 효과를 보였습니다. 특히 7B LLaVA-OV 모델과 32개의 프레임을 사용했을 때, GCR은 두 데이터셋에서 각각 64.25% 및 62.15%의 정확도를 달성하여, 가장 뛰어난 재현 성능을 보이는 기존 방법들을 각각 2.54% 및 1.93%p 앞섰습니다.

Original Abstract

Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations. First, selected frames tend to cluster around local relevance peaks, and once the budget is exhausted, omitted evidence cannot be recovered. Second, textual and visual evidence remain weakly aligned. We propose GCR, a training-free framework that casts fixed-budget frame selection as a joint evidence curation problem. Ground converts timestamped text into temporal events, selects query-relevant real frame anchors, and renders each event text onto its temporally aligned frame. Cover supplements grounded events with direct visual anchors for complementary visual evidence and applies global maximal marginal relevance to preserve diverse context. Refine revisits omitted temporal regions and replaces the weakest revisable context frame with a real-frame medoid---but only when the medoid offers greater evidence value. GCR maintains a fixed number of chronologically ordered frames and requires no VLM training or architectural modification. Experiments on LongVideoBench and Video-MME, across three 7B backbones and frame budgets of 8, 32, and 64, demonstrate consistent improvements in long-video QA. With the 7B LLaVA-OV backbone and 32 frames, GCR achieves 64.25% and 62.15% on the two benchmarks, outperforming the strongest reproduced baselines by 2.54 and 1.93 percentage points, respectively.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!