2608.05703v1 Aug 06, 2026 cs.CV

StreamArena: 지속적이고 상호작용적인 장기 예측 능력을 갖춘 에이전트 기반 스트리밍 비디오 이해를 향하여

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Shijian Wang
Shijian Wang
Citations: 89
h-index: 4
Xichen Zhang
Xichen Zhang
Citations: 72
h-index: 4
Yinghao Zhu
Yinghao Zhu
Citations: 34
h-index: 1
Sitong Wu
Sitong Wu
Citations: 231
h-index: 7
Shaozuo Yu
Shaozuo Yu
Citations: 98
h-index: 4
Meng Chu
Meng Chu
Citations: 41
h-index: 3
Jiaya Jia
Jiaya Jia
Citations: 805
h-index: 5
Yuan Lu
Yuan Lu
Citations: 44
h-index: 3
Guankai Li
Guankai Li
Citations: 30
h-index: 1

지속적인, 실제 환경에서 자율적인 다중 모달 에이전트를 운영하려면 무한한 오디오-비주얼 스트림을 처리하고 시간 단위의 장기 기억을 유지해야 합니다. 그러나 현재 평가 방법은 주로 짧은 클립과 객관식 질문 형식을 사용합니다. 이러한 설계는 마지막 4프레임만으로도 복잡한 스트리밍 모델을 능가할 수 있는 기본적인 수준의 성능을 보여주며, 또한 정답 선택지가 언어적 약어를 드러내는 문제를 야기합니다. 우리는 시간 단위의 상호작용적인 스트리밍 비디오 이해를 위한 벤치마크인 StreamArena를 소개합니다. StreamArena는 평균 88.8분의 길이로 구성된 243개의 전체 영상과, 실시간 인식, 과거 회상, 적극적인 상호작용 및 다중 모달 도구 활용을 평가하는 3,646쌍의 엄격하게 주석 처리된 개방형 질문-답변 쌍으로 구성되어 있습니다. 다양한 시스템에 대한 평가는 지속적인 상호작용과 장기적인 다중 모달 이해 사이의 긴장을 보여줍니다. 최근 프레임만 유지하는 방법은 과거의 이벤트를 복구할 수 없고, 과거 관찰 내용을 텍스트로 변환하는 방법은 시각적 증거를 잃게 되며, 시각적 기억을 반복적으로 압축하는 방법은 시간이 지남에 따라 미세한 디테일을 보존하기 어렵습니다. 우리는 이러한 긴장을 StreamMind라는 두 계층 아키텍처로 해결합니다. StreamMind는 지연 시간에 민감한 상호작용 및 적극적인 모니터링 작업을 독립적으로 예약된 프론트엔드 워커에게 할당하고, 백엔드 워커는 비동기적으로 지속적인 다중 모달 기억을 구축하고 과거 회상 및 외부 검색을 수행합니다. StreamMind는 네 가지 기능 모두에서 기존 스트리밍 모델보다 우수한 성능을 보이며, 지속적인 상태를 재사용하여 질문-답변 지연 시간을 줄입니다.

Original Abstract

Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design allows minimal baselines that process only the last four frames to match or surpass complex streaming models, while answer options also expose language shortcuts. We introduce StreamArena, a benchmark for hour-scale, interactive streaming video understanding. StreamArena contains 243 full-length videos averaging 88.8 minutes and 3,646 rigorously annotated, open-ended question-answer pairs that evaluate real-time perception, historical retrospection, proactive interaction, and multimodal tool utilization. Evaluation across diverse systems exposes a tension between continuous interaction and long-horizon multimodal comprehension. Methods that retain only recent frames cannot recover distant events, methods that convert past observations into text lose visual evidence, and methods that repeatedly compress visual memory struggle to preserve fine-grained details over time. We address this tension with StreamMind, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search. StreamMind outperforms existing streaming baselines across all four capabilities and reduces query-to-answer latency by reusing persistent state.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!