2607.13569v1 Jul 15, 2026 cs.CV

GHR-VLM: Grounded Hybrid Reasoning을 활용한 제로샷 트랜짓 비디오 분석의 실현

GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning

Ruimin Ke
Ruimin Ke
Citations: 141
h-index: 4
Kaicong Huang
Kaicong Huang
Citations: 26
h-index: 3
Weiheng Oh
Weiheng Oh
Citations: 0
h-index: 0
Jack M. Reilly
Jack M. Reilly
Citations: 4
h-index: 1
Thomas Guggisberg
Thomas Guggisberg
Citations: 1
h-index: 1

트랜짓 비디오 이해는 기존 승객 계수 및 요금 시스템으로는 파악할 수 없는 중요한 세부 정보를 제공할 수 있습니다. 그러나 지도 학습 기반 비디오 모델은 작업별 주석이 필요하며, 비전-언어 모델(VLM)을 긴 차량 내부 영상에 직접 적용하는 것은 신뢰성이 낮고 비용이 많이 듭니다. 본 연구에서는 이러한 접근 방식의 장점을 결합하기 위해, 제로샷 트랜짓 버스 비디오 분석을 위한 시각적 기반 하이브리드 추론 프레임워크인 GHR-VLM을 제안합니다. 이는 명시적인 시각적 고정(grounding)이 VLM의 추론 능력을 향상시킨다는 관찰에서 비롯되었으며, 긴 감시 영상을 압축된 승객 중심의 공간-시간 증거로 변환하여 활용합니다. 구체적으로, 경량화된 엣지 기반 모니터가 지속적으로 문 상태를 추적하고 승객 클립을 분할하는 엣지-클라우드 설계를 제안합니다. 백엔드 VLM은 이렇게 분리된 승객 클립과 관련된 데이터를 활용하여, 두 단계의 거친 수준에서 세밀한 수준으로 조정하는 과정을 통해 승차 승객을 식별하고 결제 행동을 분류합니다. GHR-VLM은 VLM이 필요한 연산을 승객 클립 및 관련 데이터에만 적용함으로써 클라우드 추론 비용을 줄이고, 특정 결제 방식에 대한 학습 데이터를 사용하지 않으며, VLM이 일반적으로 어려움을 겪는 로컬화된 증거를 제공합니다. 실제 버스 감시 영상 486분에 대한 평가 결과는 시각적 기반 엣지-클라우드 추론이 승객 단위의 결제 분석에 잠재력을 지니고 있음을 보여주며, 동시에 저품질 비디오 환경과 같은 해결해야 할 과제를 강조합니다.

Original Abstract

Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. However, supervised video models require task-specific annotations, while applying vision-language models (VLMs) directly to long onboard videos is unreliable and costly. To leverage the complementary strengths of both approaches, we propose GHR-VLM, a visual grounded hybrid reasoning framework for zero-shot transit-bus video analytics. It is motivated by the observation that explicit visual grounding can improve VLM reasoning by converting long surveillance streams into compact, passenger-centered spatiotemporal evidence. Specifically, we propose an edge-cloud design in which a lightweight edge-based monitor continuously tracks door status and segments passenger clips. A backend VLM then identifies boarding passengers and classifies payment behavior through a two-stage coarse-to-fine refinement of spatiotemporal evidence. By invoking the VLM only on grounded passenger clips and contact sheets, GHR-VLM reduces cloud inference, avoids payment-specific training data, and supplies the localized evidence that VLMs otherwise struggle to identify. Evaluation on 486 minutes of real-world bus surveillance video demonstrates the potential of grounded edge-cloud reasoning for passenger-level payment analytics while highlighting the challenges posed by degraded video conditions.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!