2607.02269v1 Jul 02, 2026 cs.CV

AnyGroundBench: 비전-언어 모델의 동영상 지칭을 위한 전문 분야 벤치마크

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Ryo Hachiuma
Ryo Hachiuma
Citations: 113
h-index: 6
Ryo Fujii
Ryo Fujii
Keio University
Citations: 208
h-index: 9
Hideo Saito
Hideo Saito
Citations: 94
h-index: 5
Rintaro Otsubo
Rintaro Otsubo
Citations: 2
h-index: 1
Reina Ishikawa
Reina Ishikawa
Citations: 61
h-index: 6
T. Kanaya
T. Kanaya
Citations: 0
h-index: 0
Kanta Sawafuji
Kanta Sawafuji
Citations: 2
h-index: 1
Hiroki Kajita
Hiroki Kajita
Citations: 26
h-index: 2
Shigeki Sakai
Shigeki Sakai
Citations: 0
h-index: 0

비전-언어 모델(VLM)은 시공간적 동영상 지칭(STVG)에서 엄청난 잠재력을 보여주었습니다. 그러나 현재의 평가 프로토콜은 대부분 일반적인 일상생활 벤치마크에 대한 제로샷 평가에 국한되어 있습니다. 이는 특정 분야의 실제 응용 분야와 중요한 단절을 초래하며, 모델이 불가피하게 희귀한 시각적 개념과 복잡한 시공간적 동역학에 직면하게 됩니다. 무한한 데이터 분포를 대상으로 하는 광범위한 사전 훈련은 실현 불가능하므로, 새로운 도메인에 대한 적응 능력은 필수적입니다. 이러한 격차를 해소하기 위해, 우리는 STVG 평가 패러다임을 정적인 제로샷 테스트에서 엄격한 도메인 적응으로 전환하도록 설계된 도메인 적응 벤치마크인 AnyGroundBench를 소개합니다. AnyGroundBench는 동물, 산업, 스포츠, 수술 및 공공 안전의 다섯 가지 전문 분야를 대상으로 하며, 전문가가 주석을 달아 만든 새로운 동영상(예: 마우스 행동)과 기존 데이터 세트를 결합하고, 밀집되고 고품질의 시공간적 주석을 통해 이를 통합합니다. 더욱 중요한 것은, 이 벤치마크는 도메인 적응성을 체계적으로 측정하기 위한 전용 학습 데이터 세트를 제공합니다. 우리는 15개의 최첨단 VLM을 광범위하게 평가하여, 실제 계산 제약 조건 하에서 이러한 모델의 제로샷 일반화 및 인컨텍스트 학습(ICL) 능력을 평가했습니다. 궁극적으로, 우리의 연구 결과는 현재 모델이 전문 분야에 직면했을 때 제로샷 및 ICL 기반 적응 모두에서 실패하며, 향후 연구가 해결해야 할 시공간적 추론의 중요한 결함을 드러낸다는 것을 보여줍니다.

Original Abstract

Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured videos such as expert-annotated mouse behaviors with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated training subsets to systematically measure domain adaptability. We extensively evaluate 15 state-of-the-art VLMs, assessing their zero-shot generalization and In-Context Learning (ICL) capabilities under practical computational constraints. Ultimately, our findings reveal that current models fail in both zero-shot and ICL-based adaptation when confronted with specialized domains, exposing critical flaws in spatio-temporal reasoning that future research must address.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!