증거 기반의 시간 정렬 및 투명한 의사 결정을 갖춘 점진적인 온라인 비디오 이해
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
야생 환경에서 작동하는 시각적 에이전트는 비디오 스트림에서 충분한 증거가 처음 나타나는 시점에 정확하게 응답해야 하며, 이는 오프라인 환경에서 평가되는 기존의 비디오 LLM에서는 간과되는 중요한 기능입니다. 온라인 스트리밍 패러다임으로의 전환은 의사 결정의 투명성 부족, 시각적 증거에 따른 응답 시간의 정렬 어려움, 그리고 제한된 계산 자원 내에서 전역적이고 인과적으로 일관된 이해를 유지해야 하는 필요성과 같은 상당한 과제를 야기합니다. 이러한 문제를 해결하기 위해, 우리는 추론 제어와 메모리 통합을 분리하는 새로운 프레임워크를 제안합니다. 우리는 이 프레임워크의 구체적인 구현체인 extbf{\model{}}을 소개합니다. 첫째, extit{Active Thinking Decision Maker (ATDM)}는 관찰 가능한 진행률 ($\boldsymbol{ρ}$) 및 신뢰도 ($\boldsymbol{c}$) 지표를 사용하여 의사 결정 과정을 외부화하는 투명한 추론 제어기입니다. 이를 통해 ATDM은 사용자가 추론 과정을 실시간으로 확인하는 동시에, 응답 시간 $t_r$을 처음으로 충분한 증거가 나타나는 시간 $t^\star$와 정확하게 일치시킬 수 있습니다. 둘째, extit{Hierarchical Progressive Semantic Integration (HPSI)} 모듈은 효율적인 메모리 시스템 역할을 합니다. HPSI는 학습 가능한 다단계 집계 토큰을 사용하여 클립 전체에 걸쳐 전파함으로써, 토큰 예산 초과 없이 풍부하고 전역적인 인지 상태를 구축합니다. % 우리 접근 방식은 주요 온라인 비디오 이해 벤치마크에서 새로운 기준을 설정하며, StreamingBench에서 71.6%, OVOBench에서 46.9%의 뛰어난 성능을 달성하여, 증거 기반의 시간 정렬 및 투명한 온라인 비디오 분석을 위한 강력한 솔루션을 입증합니다. 광범위한 실험을 통해 ATDM 및 HPSI의 효과를 입증했으며, 예를 들어 Thinking-QwenVL은 StreamingBench 벤치마크에서 이전 최고 성능인 67.63%에서 71.60%로 정확도를 향상시켰습니다.
Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challenges: a lack of decision transparency, the difficulty of aligning response timing with visual evidence, and the need to maintain a global, causally consistent understanding under tight computational budgets. To address these issues, we propose a novel framework that decouples reasoning control from memory integration. We introduce \textbf{\model{}}, an instantiation of this framework with two core components. First, the \emph{Active Thinking Decision Maker (ATDM)} is a transparent reasoning controller that externalizes its decision process using observable progress ($\boldsymbolρ$) and confidence ($\boldsymbol{c}$) metrics. This allows it to precisely time its response $t_r$ to match the first-sufficient-evidence timestamp $t^\star$ while streaming its reasoning to the user. Second, the \emph{Hierarchical Progressive Semantic Integration (HPSI)} module acts as an efficient memory system. It employs a set of learnable, multi-level aggregation tokens that are propagated across clips to build a rich, global cognitive state without exceeding token budgets. %Our approach sets a new standard on key online video understanding benchmarks, achieving strong performance of \textbf{71.6\%} on StreamingBench and \textbf{46.9\%} on OVOBench, demonstrating a robust solution for evidence-aligned and transparent online video analysis. Extensive experiments demonstrate the effectiveness of ATDM and HPSI, e.g., Thinking-QwenVL improves the accuracy of the previous state-of-the-art from 67.63\% to 71.60\% on the StreamingBench benchmark.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.