2607.13421v1 Jul 15, 2026 cs.CV

ScanFocus: 시공간 비디오 기반 객체 추적을 위한 거칠기-세밀함 프레임워크

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

Kai Chen
Kai Chen
Citations: 0
h-index: 0
Ming Dai
Ming Dai
Citations: 101
h-index: 6
Wenxuan Cheng
Wenxuan Cheng
Citations: 32
h-index: 3
Wankou Yang
Wankou Yang
Citations: 235
h-index: 4

시공간 비디오 기반 객체 추적(STVG)은 자연어 표현으로 설명된 특정 객체의 시각적 경로를 비디오 스트림에서 찾아내는 것을 목표로 합니다. 그러나 대부분의 최신 방법들은 전반적인 맥락 모델링과 정확한 경계 지역화를 균형 있게 처리하는 데 어려움을 겪습니다. 긴 비디오를 처리하는 데 드는 막대한 계산 비용 때문에, 이러한 접근 방식은 종종 낮은 프레임률의 시간 축소와 암시적 움직임 모델링을 사용합니다. 이는 불가피하게 미세한 경계 단서들을 억제하고 정확한 경계 확분을 위해 필요한 명시적인 프레임 간 의존성을 무시합니다. 이러한 한계를 극복하기 위해, 우리는 STVG 작업을 전반적인 시공간 스캔과 로컬 경계 집중으로 분리하는 새로운 거칠기-세밀함 프레임워크인 **ScanFocus**를 제안합니다. 구체적으로, 우리는 통일된 비전-언어 융합 인코더와 경량화된 변형 의미-운동 융합 모듈을 사용하여 다중 모드 특징을 효율적으로 정렬하고 거친 제안을 생성합니다. 억제된 미세한 세부 정보를 복구하기 위해, 우리는 세밀함 조정 단계에서 의미 기반 시간 집계기(SGTA)를 도입합니다. SGTA는 거친 경계를 중심으로 밀집하게 샘플링하여 의미 지침 하에 단기적인 시간적 상호 작용을 명시적으로 모델링하고, 정확한 타임스탬프 회귀를 위해 빠른 움직임 변화를 포착합니다. 세 개의 널리 사용되는 벤치마크에서 수행된 광범위한 실험은 제안하는 방법이 기존 접근 방식보다 우수한 성능을 보임을 입증했습니다. 코드는 https://github.com/TenMinutes209/ScanFocus 에서 공개될 예정입니다.

Original Abstract

Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with precise boundary localization. Due to the prohibitive computational costs of processing long videos, these approaches typically resort to low-rate temporal downsampling and implicit motion modeling. This inevitably suppresses high-frequency boundary cues and neglects the explicit inter-frame dependencies required for precise boundary delineation. To address these limitations, we present \textbf{ScanFocus}, a novel coarse-to-fine framework that decouples the STVG task into a global spatio-temporal scan and a local boundary focus. Specifically, we utilize a unified vision-language fusion encoder combined with a lightweight Deformable Semantic-Motion Fusion module to efficiently align multimodal features and generate coarse proposals. To recover the suppressed fine-grained details, we introduce the Semantic-Guided Temporal Aggregator (SGTA) in the refinement stage. By densely sampling around coarse boundaries, SGTA explicitly models short-term temporal interactions under semantic guidance, capturing rapid motion changes for precise timestamp regression. Extensive experiments on three widely used benchmarks demonstrate the performance superiority of our proposed method over previous approaches. Code will be released at https://github.com/TenMinutes209/ScanFocus.

0 Citations
0 Influential
34.989476363992 Altmetric
0.0 Score
Original PDF
10

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!