2608.03779v1 Aug 04, 2026 cs.CV

AgenticVAU: 비디오 이상 상황 이해를 위한 다중 에이전트 탐색-검증 추론

AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Yuxiang Duan
Yuxiang Duan
Citations: 3
h-index: 1
Ning Liu
Ning Liu
Citations: 30
h-index: 2
Shuai Feng
Shuai Feng
Citations: 21
h-index: 3
Lanju Kong
Lanju Kong
Citations: 0
h-index: 0
Xingdong Sheng
Xingdong Sheng
Citations: 214
h-index: 6

비디오 이상 상황 이해(Video Anomaly Understanding, VAU)는 비디오 내의 비정상적인 사건을 종합적으로 해석하는 것을 목표로 하며, 모델은 이상 징후를 식별하고, 이를 뒷받침하는 증거를 발견하며, 단순한 이상 탐지를 넘어 근본적인 원인을 설명해야 합니다. 기존의 VAU 방법들은 종종 특수한 학습이나 제한된 관찰에 의존하여 일반화 성능이나 증거 범위가 제한되는 경우가 많습니다. 단일 에이전트 방식은 적응적인 비디오 관찰을 지원하지만, 여전히 탐색, 관찰 및 의사 결정을 하나의 통합 추론 과정 내에서 수행하며, 역할 전문성과 체계적인 증거 조율 측면에서 한계를 가집니다. 이러한 제한점을 해결하기 위해, 저희는 학습이 필요 없는 다중 에이전트 프레임워크인 AgenticVAU를 제안합니다. AgenticVAU는 VAU를 탐색-검증(explore-verify) 과정으로 구성하며, 시스템은 먼저 잠재적인 이상 징후를 발견한 다음, 타겟 관찰을 통해 이를 검증합니다. 이를 위해 시각적 규칙 생성, 검색 계획 수립, 비디오 관찰 및 최종 의사 결정을 처리하는 네 가지 전문 에이전트를 도입했습니다. 이러한 에이전트들은 각 관찰 내용을 연결하는 공유 증거 메모리인 '앵커 레지스트리'를 통해 서로 통신합니다. 이 에이전트 프레임워크의 지침에 따라 AgenticVAU는 충분한 증거가 수집될 때까지 광범위한 시간적 탐색, 밀도 높은 지역 검증 및 구간 간 비교를 반복합니다. 저희는 ECVA, UCF-Crime, 그리고 VAU-Bench의 MSAD 데이터셋을 사용하여 광범위한 실험을 수행했으며, 그 결과 AgenticVAU는 사전 학습 없이 추론하는 방식과 강화 학습 기반의 기존 방법들을 능가하는 성능을 보여주었으며, 이는 다중 에이전트 협업이 비디오 이상 상황 이해에 매우 유용하다는 것을 입증합니다.

Original Abstract

Video anomaly understanding (VAU) focuses on comprehensively interpreting abnormal events in videos, requiring models to identify anomalous occurrences, discover their supporting evidence, and explain the underlying causes beyond simple anomaly detection. Existing VAU methods often rely on specialized training or limited observations, restricting generalization or evidence coverage. Although single-agent alternatives support adaptive video observation, they still integrate exploration, observation, and decision-making within a unified reasoning process, offering limited role specialization and structured evidence coordination. To address these limitations, we present AgenticVAU, a training-free multi-agent framework that casts VAU as an explore--verify process, where the system first discovers potential anomalies and then verifies them through targeted observations. To achieve this, four specialized agents are introduced to handle visual-rule construction, search planning, video observation, and final decision, respectively. These agents communicate through an anchor registry, a shared evidence memory that binds each observation. Guided by this agent framework, AgenticVAU interleaves broad temporal exploration, dense local verification, and cross-interval comparison until sufficient evidence is collected. We conduct extensive experiments on the ECVA, UCF-Crime, and MSAD subsets of VAU-Bench, the results show that AgenticVAU outperforms zero-shot inference and reinforcement learning-based baselines, demonstrating the value of multi-agent collaboration for video anomaly understanding.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!