2608.02078v1 Aug 03, 2026 cs.CL

CAVE: 비디오 시간적 위치 추정을 위한 역량 기반 시각 경계 증거 정렬

CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding

Zhicong Lu
Zhicong Lu
Citations: 63
h-index: 5
Wei Jia
Wei Jia
UCAS
Citations: 453
h-index: 3
Huaxing Liu
Huaxing Liu
Citations: 1
h-index: 1
Xiang Wang
Xiang Wang
Citations: 664
h-index: 8
Shuai Li
Shuai Li
Citations: 103
h-index: 3
Yu Chen
Yu Chen
Citations: 1
h-index: 1
Wenqiang Lv
Wenqiang Lv
Citations: 23
h-index: 2
Jiayue Cao
Jiayue Cao
Citations: 16
h-index: 1

최근, 대규모 시각-언어 모델(LVLM)은 강화 학습(RL)을 통해 비디오 시간적 위치 추정(VTG) 분야에서 상당한 성능 향상을 이루었습니다. 그러나 기존 방법들은 주로 최종 예측 구간의 정확도를 평가하는 결과 기반 보상에 의존하며, 이로 인해 경계와 관련된 시각 정보 및 해당 정보와 타임스탬프 예측 간의 연관성이 충분히 제약되지 못합니다. 본 논문에서는 타임스탬프 예측과 그 근본적인 수준인 경계 레벨의 시각 정보를 심층적으로 분석하고, 널리 사용되는 벤치마크에서 시각 정보와 예측된 타임스탬프 간의 불일치가 빈번하게 발생한다는 것을 보여줍니다. 이러한 문제를 해결하기 위해, 본 연구에서는 역량 기반 시각 경계 증거 정렬(CAVE)이라는 새로운 방법을 제안합니다. CAVE는 위치 추정 최적화에 경계별 시각 정보 보상을 추가하여 시각 정보-타임스탬프 불일치를 완화합니다. 구체적으로, CAVE는 경계별 시각 정보를 명시적으로 표현하기 위해 경계별 증거 토큰을 도입하고, 가벼운 지도 학습 방법을 통해 이들의 구조적 생성 및 뚜렷한 경계 의미를 초기화합니다. 강화 학습 과정에서, 시각 경계 증거 정렬 보상은 실제 경계 내에 존재하는 특수 증거 토큰의 시각적 주의를 강화하여 시각 정보와 시간적 경계 간의 일관성을 높입니다. 또한, 성능 기반 게이팅 메커니즘을 통해 위치 추정이 충분히 정확해진 경우에는 세밀한 경계 개선에 대한 과도한 제약을 피하면서, 추정 성능이 낮은 그룹에 대해서는 증거 지침을 지속적으로 제공합니다. 여러 공개 VTG 벤치마크에서 수행된 광범위한 실험 결과는 본 방법의 효과를 입증합니다.

Original Abstract

Large vision-language models (LVLMs) have achieved substantial performance gains in Video Temporal Grounding (VTG) through reinforcement learning (RL). However, existing methods primarily rely on outcome correctness rewards that evaluate only the final predicted intervals, leaving boundary-related visual evidence and its correspondence with timestamp predictions insufficiently constrained. In this paper, we delve into timestamp prediction and its underlying boundary-level visual evidence, showing prevalent misalignment between visual evidence and predicted timestamps across widely used benchmarks. To address this issue, we propose Competence-Aware Visual Boundary Evidence Alignment (CAVE), which augments localization optimization with boundary-specific visual evidence rewards to mitigate evidence-timestamp misalignment. Specifically, to explicitly represent the boundary-specific visual evidence, CAVE introduces boundary-specific evidence tokens and initializes their structured generation and distinct boundary semantics through a lightweight supervised warm-up. During RL, the visual boundary evidence alignment reward reinforces the visual attention of special evidence tokens within the ground-truth boundaries, thereby promoting alignment between visual evidence and temporal boundaries. Moreover, performance-aware gating for evidence supervision is designed to adaptively retain evidence guidance for poorly localized groups while reducing it once localization becomes sufficiently accurate to avoid over-constraining fine-grained boundary refinement. Extensive experiments on several public VTG benchmarks demonstrate the effectiveness of our method.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!