2608.05699v1 Aug 06, 2026 cs.CV

TAU-Bench: 이상 객체 추적에서부터 세밀한 수준의 비디오 이상 상황 이해까지

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chenxin Li
Chenxin Li
Citations: 1,132
h-index: 16
S. Xie
S. Xie
Citations: 189
h-index: 7
Kepeng Yang
Kepeng Yang
Citations: 0
h-index: 0
Rongxin Gao
Rongxin Gao
Citations: 13
h-index: 2
Zixin Su
Zixin Su
Citations: 0
h-index: 0
Rui Wu
Rui Wu
Citations: 0
h-index: 0
Panwang Pan
Panwang Pan
Citations: 498
h-index: 10
Yuzhi Huang
Yuzhi Huang
Citations: 108
h-index: 5
Yue Huang
Yue Huang
Citations: 116
h-index: 5
Jingyan Jiang
Jingyan Jiang
Citations: 145
h-index: 6

사람은 비정상적인 사건을 일관된 인지 과정을 통해 이해하는데, 이는 중심 객체를 식별하고 그 객체의 행동 변화를 따라가는 과정과 함께 주변 환경과의 불일치를 해석하는 것을 포함합니다. 비디오 이상 상황 이해(VAU)는 모델에게 이러한 능력을 부여하고자 하며, 단순히 비디오가 이상적인지 여부를 판단하는 것에서 벗어나 사건의 전개 방식과 중요성을 설명하는 방향으로 나아가고자 합니다. 최근의 시각-언어 모델(VLM)은 상세하고 설득력 있는 이상 상황 설명을 생성할 수 있지만, 이러한 해석이 시간이 지남에 따라 정확한 이상 객체와 일치하는지 보장하지 않습니다. 기존 벤치마크는 일반적으로 추적 및 의미 이해를 별도의 절차로 평가하여, 객체-의미 불일치를 제대로 측정하지 못합니다. 따라서 우리는 객체 추적 중심의 벤치마크인 TAU-Bench를 소개하며, 이를 통해 비디오 이상 객체 추적과 세밀한 수준의 이상 상황 이해를 동시에 평가할 수 있습니다. TAU-Bench는 49개의 사건 범주와 45개의 장면 범주에 걸쳐 1,118개의 비디오, 1,454개의 추적 정보, 그리고 202,438개의 픽셀 단위 마스크를 포함하며, 객체 수준 식별, 사건 수준 이해, 그리고 장면 수준 추론을 연결하는 객체 중심의 주석이 함께 제공됩니다. TAU-Bench를 대규모로 구축하기 위해 우리는 이상 상황 적합성 필터링, 이상 객체 추적 생성, 계층적 캡션 어노테이션, 그리고 인간 품질 관리를 통합한 자동화된 데이터 엔진을 개발했습니다. 대표적인 VLM 패밀리에 대한 평가 결과는, 그럴듯한 이상 상황 해석을 생성하는 모델조차도 정확한 객체를 안정적으로 찾고 추적하는 데 실패할 수 있으며, 이는 의미 추론과 시각적 기반 사이의 지속적인 격차를 보여줍니다. 이러한 결과는 신뢰성 있는 VAU 시스템 개발을 위한 중요한 단계로, 객체 기반 평가의 중요성을 강조합니다.

Original Abstract

Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision--language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance--semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!