오디오-비디오 엔티티 결합 및 에이전트 기반 검색을 이용한 계층적 장편 비디오 이해
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
장편 비디오 이해는 매우 긴 문맥 정보를 처리해야 하므로, 시각-언어 모델에게 상당한 어려움을 안겨줍니다. 기존의 단순 분할 전략과 검색 기반 생성 방식을 사용하는 방법은 종종 정보 파편화 및 전체적인 일관성 부족 문제를 야기합니다. 본 논문에서는 오디오-비디오 엔티티 결합과 계층적 비디오 인덱싱을 에이전트 기반 검색과 통합하여 일관성 있고 포괄적인 추론을 가능하게 하는 통합 프레임워크인 HAVEN을 제시합니다. 먼저, 시각 및 청각 스트림에서 엔티티 수준의 표현을 통합하여 의미적 일관성을 유지하고, 전체 요약, 장면, 세그먼트 및 엔티티 레벨을 아우르는 구조화된 계층으로 콘텐츠를 구성합니다. 그런 다음, 에이전트 기반 검색 메커니즘을 사용하여 이러한 계층에서 동적으로 정보를 검색하고 추론함으로써 일관성 있는 내러티브 재구성과 세밀한 엔티티 추적을 지원합니다. 광범위한 실험 결과, 제안하는 방법은 우수한 시간적 일관성, 엔티티 일관성 및 검색 효율성을 달성하며, LVBench 데이터셋에서 84.1%의 전반적인 정확도를 기록하여 새로운 최고 성능을 달성했습니다. 특히, 어려운 추론 영역에서 80.1%라는 뛰어난 성능을 보였습니다. 이러한 결과는 구조화된 다중 모드 추론이 장편 비디오의 포괄적이고 문맥 일관적인 이해에 얼마나 효과적인지를 보여줍니다.
Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We present HAVEN, a unified framework for long-video understanding that enables coherent and comprehensive reasoning by integrating audiovisual entity cohesion and hierarchical video indexing with agentic search. First, we preserve semantic consistency by integrating entity-level representations across visual and auditory streams, while organizing content into a structured hierarchy spanning global summary, scene, segment, and entity levels. Then we employ an agentic search mechanism to enable dynamic retrieval and reasoning across these layers, facilitating coherent narrative reconstruction and fine-grained entity tracking. Extensive experiments demonstrate that our method achieves good temporal coherence, entity consistency, and retrieval efficiency, establishing a new state-of-the-art with an overall accuracy of 84.1% on LVBench. Notably, it achieves outstanding performance in the challenging reasoning category, reaching 80.1%. These results highlight the effectiveness of structured, multimodal reasoning for comprehensive and context-consistent understanding of long-form videos.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.