2603.17307v1 Mar 18, 2026 cs.CV

심포니: 인지 기반의 다중 에이전트 시스템을 활용한 장편 비디오 이해

Symphony: A Cognitively-Inspired Multi-Agent System for Long-Video Understanding

Haiyang Yan
Haiyang Yan
Citations: 62
h-index: 5
Peng Xu
Peng Xu
Citations: 10,019
h-index: 17
Xiao Feng
Xiao Feng
Citations: 14
h-index: 2
Mengyi Liu
Mengyi Liu
Citations: 16
h-index: 2
Hongyu Zhou
Hongyu Zhou
Citations: 2
h-index: 1

MLLM 에이전트의 빠른 발전과 광범위한 적용에도 불구하고, 여전히 높은 정보 밀도와 긴 시간 범위를 특징으로 하는 장편 비디오 이해(LVU) 과제에서 어려움을 겪고 있습니다. 최근 LVU 에이전트에 대한 연구에서는 단순한 작업 분해 및 협업 메커니즘이 장기간의 추론 작업에는 충분하지 않다는 점이 밝혀졌습니다. 또한, 임베딩 기반 검색을 통해 직접적으로 시간 범위를 줄이는 것은 복잡한 문제의 중요한 정보를 잃게 만들 수 있습니다. 본 논문에서는 이러한 한계를 극복하기 위해 다중 에이전트 시스템인 '심포니'를 제안합니다. 심포니는 인간의 인지 패턴을 모방하여 LVU를 세분화된 하위 작업으로 분해하고, 반성을 통해 강화된 심층적인 추론 협업 메커니즘을 통합하여 추론 능력을 향상시킵니다. 또한, 심포니는 VLM 기반의 접지 방식을 제공하여 LVU 과제를 분석하고 비디오 세그먼트의 관련성을 평가함으로써, 암묵적인 의도와 큰 시간 범위를 가진 복잡한 문제를 찾아내는 능력을 크게 향상시킵니다. 실험 결과, 심포니는 LVBench, LongVideoBench, VideoMME, 및 MLVU에서 최첨단 성능을 달성했으며, LVBench에서 기존 최고 성능 모델보다 5.0% 향상된 결과를 보였습니다. 코드 및 관련 자료는 https://github.com/Haiyang0226/Symphony 에서 확인할 수 있습니다.

Original Abstract

Despite rapid developments and widespread applications of MLLM agents, they still struggle with long-form video understanding (LVU) tasks, which are characterized by high information density and extended temporal spans. Recent research on LVU agents demonstrates that simple task decomposition and collaboration mechanisms are insufficient for long-chain reasoning tasks. Moreover, directly reducing the time context through embedding-based retrieval may lose key information of complex problems. In this paper, we propose Symphony, a multi-agent system, to alleviate these limitations. By emulating human cognition patterns, Symphony decomposes LVU into fine-grained subtasks and incorporates a deep reasoning collaboration mechanism enhanced by reflection, effectively improving the reasoning capability. Additionally, Symphony provides a VLM-based grounding approach to analyze LVU tasks and assess the relevance of video segments, which significantly enhances the ability to locate complex problems with implicit intentions and large temporal spans. Experimental results show that Symphony achieves state-of-the-art performance on LVBench, LongVideoBench, VideoMME, and MLVU, with a 5.0% improvement over the prior state-of-the-art method on LVBench. Code is available at https://github.com/Haiyang0226/Symphony.

1 Citations
0 Influential
38.897207708399 Altmetric
6.9 Score
Original PDF
7

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!