2605.29402v1 May 28, 2026 cs.CV

효율적인 장편 비디오 추론을 위한 의미 및 시각적 증거: HD-EPIC VQA 챌린지를 위한 솔루션

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Liuxin Zhang
Liuxin Zhang
Citations: 260
h-index: 10
Yinsong Xu
Yinsong Xu
Citations: 84
h-index: 5
Wei Jing
Wei Jing
Citations: 931
h-index: 10
Wanjun Lv
Wanjun Lv
Citations: 7
h-index: 1
Hui Li
Hui Li
Citations: 19
h-index: 2

멀티모달 대규모 언어 모델(MLLM)은 제한된 컨텍스트 길이와 세밀한 시각적 정보의 부족으로 인해 장편 일인칭 비디오를 이해하는 데 어려움을 겪습니다. 최근에 제안된 HD-EPIC 벤치마크는 이러한 한계를 강조하며, 강력한 장기 컨텍스트 모델조차도 다양한 비디오 질문 응답 작업에서 상대적으로 낮은 성능을 보입니다. 본 논문에서는 장편 비디오 추론을 상호 보완적인 두 가지 형태의 증거, 즉 의미적 증거와 시각적 증거로 분리하는 통합 프레임워크를 제안합니다. 의미적 증거는 조잡화된 방식에서 세밀하게 추출하여 전체적인 절차 구조를 파악하며, 객체 중심적인 시각적 증거는 경계 상자와 시각적 임베딩을 통해 세밀한 정보의 연결성을 유지합니다. 추론 과정에서 우리는 추론을 쿼리에 조건부로 설정된 증거 검색 및 통합 프로세스로 정의하고, 관련 정보를 양쪽 소스에서 동적으로 선택합니다. 제안하는 방법은 HD-EPIC-VQA 챌린지의 여러 작업 범주에서 경쟁력 있는 성능을 달성했습니다. 더 넓게 보면, 본 연구 결과는 MLLM을 사용한 효과적인 장편 비디오 이해를 위해서는 의미적 및 시각적 증거를 명시적으로 구조화하고 검색하며 통합하는 것이 중요하다는 것을 보여줍니다.

Original Abstract

Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insufficient grounding of fine-grained visual details. The recently proposed HD-EPIC benchmark highlights these limitations: even strong long-context models achieve relatively low performance across diverse video question answering tasks. In this paper, we propose a unified framework that decouples long-video reasoning into two complementary forms of evidence: semantic evidence and visual evidence. Semantic evidence captures global procedural structure through a coarse-to-fine extraction pipeline, while object-centric visual evidence preserves fine-grained grounding through bounding boxes and visual embeddings. During inference, we formulate reasoning as a query-conditioned evidence retrieval and integration process, dynamically selecting relevant information from both sources. Our approach achieves competitive performance in the HD-EPIC-VQA Challenge across multiple task categories. More broadly, our results demonstrate that explicitly structuring, retrieving, and integrating semantic and visual evidence is critical for effective long-video understanding with MLLMs.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!