2607.23951v1 Jul 27, 2026 cs.CV

TimePLE: 비디오 시간적 위치 파악을 위한 시간 표현 방식 재고

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Jinfa Huang
Jinfa Huang
Citations: 960
h-index: 15
Yuhui Zeng
Yuhui Zeng
Citations: 14
h-index: 2
Xiawu Zheng
Xiawu Zheng
Citations: 27
h-index: 3
Xinyu Mao
Xinyu Mao
Citations: 20
h-index: 2
Xiaokun Liu
Xiaokun Liu
Citations: 256
h-index: 4
Xin Tao
Xin Tao
Citations: 58
h-index: 4
Jiayi Ji
Jiayi Ji
Citations: 9
h-index: 2

비디오 시간적 위치 파악(VTG)은 자연어 질의에 의해 설명되는 연속적인 비디오 구간을 특정하는 것을 목표로 합니다. 그러나 현재 대부분의 VLM 기반 방법은 이 구간을 두 개의 출력 지점을 통해 간접적으로 생성하는데, 이는 일반적으로 개별 타임스탬프 토큰 또는 연속적인 경계 좌표로 표현됩니다. 이러한 방식들은 출력 지점의 인코딩 방식에 차이가 있지만, 예측하는 대상에는 차이가 없습니다. 즉, 이벤트 구간은 여전히 파생된 객체이며, 구간의 유효성, 지속 시간 및 구간 수준의 유사성은 암묵적으로만 처리됩니다. 본 논문에서는 TimePLE을 제안합니다. TimePLE은 VTG를 출력 지점 예측에서 벗어나 구간 자체에 기반한 방식으로 재정의하며, 유효한 시간적 구간에 대한 단일의 결합 분포를 예측합니다. TimePLE은 각 구간을 표준 위치-지속 시간 사각형 내의 한 점으로 매핑하며, 여기서 각 지점은 유효한 구간에 해당하고 인접한 점들은 기하학적으로 유사한 구간을 나타냅니다. 비디오와 질의가 주어지면, VLM은 단일의 잠재적인 <|TIMESPAN|> 토큰을 생성하고, 이 토큰의 숨겨진 상태는 결합된 구간 분포로 디코딩됩니다. 이 분포는 지속 시간 인식을 기반으로 한 좌표 보정 과정을 거쳐 연속적인 경계 값으로 변환됩니다. 동일한 구간 표현은 입력 시간 기준점을 인코딩하는 데 사용되어, 비디오 측면의 시간적 증거와 출력 측면의 구간 예측을 연결합니다. 잠재적인 구간 표현과 완전한 이벤트 구간을 안정적으로 정렬하기 위해, 9만 규모의 데이터셋을 구축하고 3천 규모의 벤치마크 주석에 대해 인간 검증을 수행했습니다. 네 가지 VTG 벤치마크를 사용한 실험 결과, TimePLE은 출력 지점 예측 기반의 기존 방법들을 지속적으로 능가하며, 평균 mIoU 점수가 58.9%로 나타났습니다. 특히 짧은 지속 시간 및 중간 지속 시간 이벤트에서 상당한 성능 향상을 보였습니다.

Original Abstract

Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods typically produce this interval indirectly through two endpoint outputs, represented either as discrete timestamp tokens or continuous boundary coordinates. These formulations differ in how endpoints are encoded, but not in what is predicted: the event interval remains a derived object, while interval validity, duration, and interval-level similarity are handled only implicitly. We propose TimePLE, which reformulates VTG from endpoint prediction to interval-native grounding by predicting a single joint distribution over valid temporal intervals. TimePLE maps each interval to a point in a canonical position-duration square, where every support point corresponds to a valid span and neighboring points represent geometrically similar intervals. Given a video and query, the VLM generates a single latent <|TIMESPAN|> token whose hidden state is decoded into a joint interval distribution, refined through duration-aware coordinate correction, and converted into continuous boundaries. The same interval representation is used to encode input temporal anchors, aligning video-side temporal evidence with output-side span prediction. To reliably align the latent span representation with complete event intervals, we curate 90K-scale grounded samples and human-verify 3K-scale benchmark annotations. Experiments across four VTG benchmarks show that TimePLE consistently outperforms endpoint prediction baselines, achieving an average mIoU of 58.9, with clear gains on short-duration and medium-duration events.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!