2607.28312v1 Jul 30, 2026 cs.CV

ObjectStream: 잠재 객체를 스트리밍 비디오 이해를 위한 메모리 고정점으로 활용

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

Xu Zheng
Xu Zheng
Citations: 432
h-index: 13
Mingkang Dong
Mingkang Dong
Citations: 6
h-index: 2
Bin Ren
Bin Ren
Citations: 68
h-index: 4
Mohamed Elhoseiny
Mohamed Elhoseiny
Citations: 30
h-index: 2
Muxin Pu
Muxin Pu
Citations: 2
h-index: 1
Jie Li
Jie Li
Citations: 34
h-index: 3
Bohan Guo
Bohan Guo
Citations: 0
h-index: 0
Song Chen
Song Chen
Citations: 0
h-index: 0
Chenlan Zhao
Chenlan Zhao
Citations: 0
h-index: 0
Tianwen Qian
Tianwen Qian
Citations: 557
h-index: 9
Yuqian Fu
Yuqian Fu
Citations: 368
h-index: 10

스트리밍 비디오 이해는 모델이 향후 질문이 주어지기 전에 유용한 시각적 증거를 지속적으로 유지해야 하는 것을 요구합니다. 기존 접근 방식은 주로 토큰 중요도, 시간 중복 또는 세그먼트 수준의 관련성을 기준으로 증가하는 시각적 컨텍스트를 관리하지만, 시간이 지남에 따라 지속되고 진화하는 객체 중심으로 증거를 구성하는 경우는 드뭅니다. 따라서 본 논문에서는 스트리밍 비디오 이해를 위한 메모리 고정점으로 잠재 객체를 활용하는 학습이 필요 없는 프레임워크인 ObjectStream을 소개합니다. ObjectStream은 동결된 Video-LLM 표현에서 직접 공간적으로 일관된 잠재 객체를 유도하고, 이를 프레임 간에 연결하여 지속적인 고정점을 만들고, 외부 객체 감지기나 분할 모델 없이 제한된 메모리 예산 내에서 해당 객체의 히스토리를 유지합니다. 이러한 고정점을 기반으로 ObjectStream은 지속적인 객체 히스토리, 일시적인 객체 변화 및 최근 시각적 컨텍스트의 세 가지 상호 보완적인 증거를 보존합니다. 이 설계는 기존 Video Large Language Models (Video-LLMs)가 객체의 동일성, 상호 작용 및 상태 변화에 대해 추론할 수 있도록 하면서 기본 모델을 변경하지 않습니다. 온라인 스트리밍 및 오프라인 장편 비디오 벤치마크에서의 광범위한 실험 결과는 효과성과 효율성을 모두 입증합니다. 온라인 스트리밍 평가에서 ObjectStream은 OVO-Bench Real-Time Visual Perception에서 Qwen2.5-VL-7B의 성능을 10.0점 향상시키면서, GPU 메모리 사용량 및 TTFT를 약 50% 줄였습니다. 오프라인 장편 비디오 벤치마크에서는 전체 토큰 기준을 능가하는 동시에 시각적 토큰의 82.5%를 제거했습니다. 이러한 결과는 잠재 객체가 간결한 스트리밍 비디오 메모리를 위한 실용적이고 효과적인 구성 원칙임을 강조합니다.

Original Abstract

Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understanding. ObjectStream induces spatially coherent latent objects directly from frozen Video-LLM representations, links them across frames into persistent anchors, and maintains their histories under a bounded memory budget, without requiring external object detectors or segmentation models. Built on these anchors, ObjectStream preserves three complementary forms of evidence: persistent object histories, transient object changes, and recent visual context. This design enables existing Video Large Language Models (Video-LLMs) to reason over object identities, interactions, and state changes while leaving the underlying model unchanged. Extensive experiments on online streaming and offline long-video benchmarks demonstrate both effectiveness and efficiency. In online streaming evaluation, ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, while reducing peak GPU mem-ory and TTFT by approximately 50%. On offline long-video benchmarks, it surpasses the full-token baseline while discarding 82.5% of visual tokens. These results highlight latent objects as a practical and effective organizing principle for compact streaming video memory.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!