더 많이 보고, 더 깊이 생각하기: 검색 확장 기반 시각적 증거와 답변-단서 가이드 기반 숙고를 통한 장편 비디오 이해
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding
최근 비디오 대규모 언어 모델(Video-LLM)의 발전은 장편 비디오 이해 작업에서 뛰어난 성능을 가능하게 했습니다. 그러나 기존 방법은 여전히 두 가지 주요 한계점을 가지고 있습니다. 증거 획득은 종종 단일 검색 의도에 의존하며, 답변 생성에는 효과적인 시각적 피드백 메커니즘이 부족합니다. 이러한 제한 사항을 해결하기 위해, 장편 비디오 이해를 위한 종합적인 시각적 증거 및 숙고 프레임워크인 **CoVER**를 제안합니다. CoVER는 Video-LLM이 **더 많이 보기(See More)** 위해 동적으로 검색 확장 기반의 시각적 증거를 수집하고, **더 깊이 생각하기(Think Deeper)** 위해 효과적인 답변별 시각적 피드백을 통해 초안 답변을 검증하도록 합니다. 이러한 메커니즘은 장편 비디오 이해를 답변 중심의 생성에서 증거 중심적이고 시각적으로 검증 가능한 추론으로 전환합니다. 실험 결과, CoVER-7B는 동일한 파라미터 규모를 가진 모델보다 훨씬 뛰어난 성능을 보이며, 특정 지표에서는 최첨단 상용 모델조차 능가하는 것으로 나타났습니다.
Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods still face two key limitations: evidence acquisition often relies on a single search intent, and answer generation lacks an effective visual feedback mechanism. To address these limitations, we propose \textbf{CoVER}, a Comprehensive Visual Evidence and Reflection framework for long-video understanding. CoVER enables Video-LLMs to \textbf{See More} by dynamically gathering query-expanded visual evidence, and \textbf{Think Deeper} by verifying draft answers with effective answer-specific visual feedback. Together, these mechanisms shift long-video understanding from answer-centric generation to evidence-centric and visually verifiable reasoning. Experimental results show that CoVER-7B substantially outperforms models with the same parameter scale and even surpasses state-of-the-art closed-source models on certain metrics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.