2604.23724v1 Apr 26, 2026 cs.CV

확대하여 분석, 논리적으로 판단: 베이지안 추론 기반의 집중형 VLM 추론을 통한 고속도로 감시 영상에서의 효율적인 원거리 이상 감지

Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference

Xiaowei Mao
Xiaowei Mao
Citations: 186
h-index: 8
S. Guo
S. Guo
Citations: 6,003
h-index: 18
Bowen Sui
Bowen Sui
Citations: 42
h-index: 2
Weijie Zhang
Weijie Zhang
Citations: 19
h-index: 2
Yawen Yang
Yawen Yang
Citations: 20
h-index: 2
Shi-Shun Zhao
Shi-Shun Zhao
Citations: 0
h-index: 0
Jiaqi Lin
Jiaqi Lin
Citations: 2
h-index: 1
Tin-Tu Wu
Tin-Tu Wu
Citations: 2
h-index: 1
Youfang Lin
Youfang Lin
Citations: 600
h-index: 14
Huaiyu Wa
Huaiyu Wa
Citations: 0
h-index: 0

고속도로 영상 이상 감지는 안전 관리에 필수적입니다. 그러나 다양한 장면에서 이상을 식별하는 것은 여전히 어렵고, 특히 미묘한 비정상적인 차량 움직임을 보이는 원거리 대상에 대한 감지가 더욱 어렵습니다. 비전-언어 모델(VLM)은 강력한 의미 추론 능력을 보여주지만, 전체 프레임을 처리하면 이러한 원거리 객체에 대한 주의 집중도가 낮아지고 계산 비용이 매우 높아집니다. 이러한 문제점을 해결하기 위해, 우리는 베이지안 추론에 의해 안내되는 VLM을 활용하는 비동기 협업 프레임워크인 VIBES를 제안합니다. 특히, 다양한 고속도로 환경에서의 일반화 성능 저하 문제를 해결하기 위해, 온라인 베이지안 추론 모듈을 도입했습니다. 이 모듈은 지속적으로 차량 궤적을 평가하여 정상적인 주행 행동의 확률적 경계를 동적으로 업데이트하며, 이를 비동기 트리거로 사용하여 공간 및 시간적으로 이상을 정확하게 위치시킵니다. VLM은 전체 비디오 스트림을 처리하는 대신, 트리거에 의해 지정된 특정 시각적 영역만 처리합니다. 이러한 표적 시각 입력은 주의 집중도를 높이고 정확한 의미 추론을 가능하게 합니다. 광범위한 실험 결과는 VIBES가 원거리 이상 감지 정확도를 향상시키고 계산 오버헤드를 줄이며, 다양한 고속도로 환경에서 높은 실시간 효율성과 설명 가능성을 달성함을 보여줍니다.

Original Abstract

Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-field vehicles with subtle abnormal motion. Vision-Language Models (VLMs) provide strong semantic reasoning capabilities, yet processing full frames can dilute evidence from distant targets and introduce substantial computational overhead. To address these challenges, we propose VIBES, an asynchronous framework that uses Bayesian inference to guide focused VLM reasoning. Specifically, an online kinematics-guided Bayesian inference module continuously estimates a context-dependent normal-motion distribution from vehicle trajectories and updates its probabilistic boundaries. Deviations from these boundaries produce asynchronous triggers that localize candidate anomalies in time and space. Instead of processing continuous full-frame video, the VLM reasons only over selected frames and localized visual regions associated with the triggers, reducing irrelevant visual content and unnecessary inference. Extensive experiments show that VIBES improves far-field anomaly detection and semantic interpretation while achieving real-time processing efficiency across diverse expressway conditions.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!