2605.28604v1 May 27, 2026 cs.CV

비디오 내 중요 인물 식별을 위한 다중 모드 시공간 정보 활용

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification

Wenke Huang
Wenke Huang
Citations: 1,958
h-index: 19
Bin Yang
Bin Yang
Citations: 27
h-index: 3
Xiao Wang
Xiao Wang
Citations: 404
h-index: 10
Minglei Yang
Minglei Yang
Citations: 14
h-index: 2
Zheng Wang
Zheng Wang
Citations: 29
h-index: 3
Xin Xu
Xin Xu
Citations: 13
h-index: 3
Mang Ye
Mang Ye
Citations: 20
h-index: 2

자동 비디오 편집 및 지능형 감시 시스템과 같은 응용 분야에서 비디오 장면 속의 주요 인물을 식별하는 것은 필수적입니다. 현재 대부분의 방법은 정적인 이미지와 즉각적인 시각적 단서에 주로 집중하며, 비디오에 내재된 풍부한 시공간 정보를 간과합니다. 이는 시간적 중요도 변화(Temporal Importance Shift, TIS)라는 현상을 야기하는데, 초기 프레임에서 중요한 것으로 판단되는 인물이 전체적인 시간 맥락을 고려할 때 중요도가 낮아질 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 텍스트 기반 설명과 함께 비디오 내에서 가장 영향력 있는 인물을 자동으로 식별하는 '비디오 중요 인물(Video Important Person, VIP) 식별'이라는 작업을 제안합니다. 우리는 9,249개의 비디오 세그먼트를 포함하고 있으며, 11개 범주에 걸쳐 정렬된 중요도 설명을 제공하는 대규모 데이터셋인 Temporal-VIP를 구성했습니다. TIS 문제를 완화하기 위해, 우리는 다중 모드 시공간 정보를 추출하는 소셜 큐 인코더(Social Cue Encoder, SCE), 계층적 큐 통합 및 교차 모드 정렬을 위한 시간적 중요도 보정기(Temporal Importance Rectifier, TIR), 그리고 인물 순위를 결정하는 VIP 추론 모델을 포함하는 VIP-Net 프레임워크를 개발했습니다. 실험 결과는 VIP-Net이 67.3%의 정확도를 달성하여 최첨단 모델(37.5%-53.9%)보다 현저히 우수한 성능을 보이며, 특징 기반 LLM 정제를 통해 기준 진실과 평균 0.63의 유사성을 나타냄을 보여줍니다. 데이터셋 및 코드는 https://huggingface.co/datasets/yml2002/Temporal-VIP 에서 이용 가능합니다.

Original Abstract

Identifying key individuals in video scenes is essential for applications such as automated video editing and intelligent surveillance. Current methods primarily focus on static images and immediate visual cues, overlooking the rich spatio-temporal information in videos. This leads to the phenomenon of Temporal Importance Shift (TIS), wherein individuals deemed significant in early frames may be demoted as the entire temporal context is considered. To address this, we introduce the Video Important Person (VIP) identification task, aimed at automatically identifying the most influential individuals in videos while providing textual rationales. We present Temporal-VIP, a large-scale rationale-annotated dataset consisting of 9,249 video segments across 11 categories with aligned importance rationales. To mitigate TIS, we develop the VIP-Net framework, which includes a Social Cue Encoder (SCE) for extracting multi-modal spatio-temporal cues, a Temporal Importance Rectifier (TIR) for hierarchical cue fusion and cross-modal alignment, and VIP Inference for ranking individuals. Experimental results show that VIP-Net achieves 67.3% accuracy, significantly outperforming state-of-the-art models (37.5%-53.9%) and yielding a mean rationale similarity of 0.63 to ground truth through feature-guided LLM refinement. The dataset and code are available at https://huggingface.co/datasets/yml2002/Temporal-VIP.

0 Citations
0 Influential
29.5 Altmetric
147.5 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!