원문 기반 다중 구간 검색을 위한 연구: One-to-Many Temporal Grounding
Towards One-to-Many Temporal Grounding
시간적 구간 정합(Temporal Grounding, TG)은 텍스트 질의에 해당하는 비디오 세그먼트를 특정하는 것을 목표로 합니다. 기존 연구는 주로 단일 세그먼트 검색에 초점을 맞추었습니다. 그러나 실제 시나리오에서는 종종 하나의 질의에 대해 여러 개의 서로 다른 세그먼트를 찾아야 하는 경우가 많습니다. 이러한 상황을 우리는 One-to-Many Temporal Grounding (OMTG)이라고 정의합니다. 기존 최고 성능의 다중 모드 대규모 언어 모델(MLLM)은 주로 일대일 설정에 최적화되어 있어, 이벤트 개수 인식이 부족하여 OMTG 환경에서는 종종 매우 낮은 점수를 보입니다. 이러한 격차를 해소하기 위해, 우리는 세 가지 주요 기여를 포함하는 체계적인 솔루션을 제시합니다. 첫째, Count Accuracy (C-Acc) 및 Effective Temporal F1 (EtF1)을 평가 지표로 사용하여 OMTG에 대한 최초의 포괄적인 벤치마크를 구축했습니다. 둘째, 정교한 구성 파이프라인을 통해 고품질의 56,000개 샘플로 구성된 OMTG 데이터 세트를 구축했습니다. 셋째, OMTG에 특화된 새로운 시간적 보상 함수와 캡션 보상 함수를 개발했습니다. 특히, 캡션 보상은 밀집된 비디오 캡션을 활용한 Chain-of-Thought 추론을 통해 정책 최적화를 정확성과 완전성 모두 향하도록 명시적으로 유도합니다. 광범위한 실험 결과, 제안하는 모델은 OMTG 벤치마크에서 43.65%의 새로운 최고 성능인 EtF1 점수를 달성했으며, 이는 Gemini 2.5 Pro 및 Seed-1.8보다 각각 15.85% 및 15.61% 더 높은 수치입니다.
Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, however, often require localizing multiple disjoint segments for a single query -- a setting we term One-to-Many Temporal Grounding (OMTG). Previous state-of-the-art MLLMs, optimized for one-to-one settings, struggle in this context, often yielding near-zero scores due to a lack of event cardinality perception. To bridge this gap, we present a systematic solution with three key contributions. First, we establish the first comprehensive OMTG benchmark, introducing Count Accuracy (C-Acc) and Effective Temporal F1 (EtF1) as evaluation metrics. Second, we curate a high-quality OMTG dataset comprising 56k samples through a sophisticated construction pipeline. Third, we develop novel temporal and caption reward functions specifically designed for OMTG. In particular, the caption reward leverages Chain-of-Thought reasoning over dense video captions to explicitly guide policy optimization toward both preciseness and completeness. Extensive experiments show our model achieves a new state-of-the-art EtF1 of 43.65\% on OMTG Bench, outperforming Gemini 2.5 Pro and Seed-1.8 by 15.85\% and 15.61\%, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.