프레임 선택을 넘어: 장편 비디오 이해를 위한 생성적 잠재 증거 집계
Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
장편 비디오 이해는 일반적으로 비디오를 답변 생성을 위해 작은 프레임 또는 시각적 토큰 세트로 압축합니다. 기존의 효율적인 파이프라인은 관련 시각적 내용을 명시적인 증거로 유지하는 데 중점을 둡니다. 그러나 증거를 제공하는 것만으로는 답변을 위해 서로 다른 순간에 걸쳐 상호 보완적인 정보를 통합하도록 보장할 수 없습니다. 우리의 핵심 아이디어는 생성 전에 선택된 프레임을 질문과 관련된 교차 프레임 증거로 구성하는 것입니다. 우리는 이 사후 선택 단계를 잠재 증거 인터페이스로 정의하고, 이를 GenEvA (Generative Latent Evidence Aggregation)라는 분포 기반의 잠재 증거 집계 프레임워크로 구현했습니다. 구체적으로, GenEvA는 질문에 조건화된 증거 분포를 사용하여 관련 프레임에 집중하여 집계를 수행하며, 각 프레임의 정보를 바탕으로 간결한 교차 프레임 잠재 증거를 형성합니다. 교차 프레임 통합이 항상 필요한 것은 아니므로, 동일한 분포가 이 잠재적 보완을 삽입할지 여부를 결정합니다. 네 가지 벤치마크와 두 개의 Video-MLLM 백본에서 GenEvA는 기존 방식보다 일관되게 성능 향상을 보여주었습니다. 8개의 프레임을 사용할 때, GenEvA는 LLaVA-Video의 평균 정확도를 +5.2 포인트, Qwen2.5-VL의 LVBench 정확도를 +10.1 포인트 향상시켰습니다. 이러한 성능 향상은 평균적으로 비디오 토큰 오버헤드가 0.11%에서 0.40%에 불과하며, 추가 분석 결과는 작업 인식 할당 및 Adaptive Evidence Invocation을 통한 이점을 확인했습니다.
Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answering. Our key idea is to organize selected frames into query-relevant cross-frame evidence before generation. We formulate this post-selection stage as a latent evidence interface and instantiate it with GenEvA ($\textbf{Gen}erative$ $Latent$ $\textbf{Ev}idence$ $\textbf{A}ggregation$), a distribution-guided latent evidence aggregation framework. Specifically, GenEvA uses a query-conditioned evidence distribution to focus aggregation on relevant frames, forming compact cross-frame latent evidence from their frame-specific information. Since cross-frame integration is not always needed, the same distribution determines whether to insert this latent complement. Across four benchmarks and two Video-MLLM backbones, GenEvA consistently improves matched-frame baselines. At 8 frames, it raises the four-benchmark LLaVA-Video average by $+5.2$ points and Qwen2.5-VL accuracy on LVBench by $+10.1$ points. These gains require only $0.11\%$--$0.40\%$ average video-token overhead; analyses further show task-aware allocation and benefits from Adaptive Evidence Invocation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.