2608.03269v1 Aug 04, 2026 cs.CV

클러스터 기반 프로토타입 블렌딩을 통한 효율적인 비디오 데이터 증류

Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending

Guang Li
Guang Li
Citations: 490
h-index: 12
Takahiro Ogawa
Takahiro Ogawa
Citations: 2,559
h-index: 22
Chongle Ren
Chongle Ren
Citations: 0
h-index: 0
Wenbo Huang
Wenbo Huang
Citations: 778
h-index: 11
Naoki Saito
Naoki Saito
Citations: 56
h-index: 4
Miki Haseyama
Miki Haseyama
Citations: 40
h-index: 2

비디오 데이터 증류는 대규모 비디오 데이터셋을 압축하여 학습 효과를 유지하는 소규모의 대체 데이터셋을 만드는 것을 목표로 합니다. 기존 방법들은 반복적인 최적화를 통해 압축된 비디오를 생성하는데, 이는 시간 차원 때문에 계산 비용이 증가합니다. 본 연구에서는 저장된 비디오에 대한 기울기 기반 최적화 없이도 효율적인 증류 비디오를 구성할 수 있는지 조사합니다. 이러한 구성 기반 접근 방식은 세 가지 과제를 해결해야 합니다: 정보가 풍부한 시각적 부분을 선택하는 것, 제한된 클래스당 비디오 수를 고려하여 다양한 클래스 내 변형을 포괄하는 것, 그리고 각 저장된 샘플이 포함하는 정보를 증가시키는 것입니다. 이를 위해 효율적인 선택-할당-블렌딩 프레임워크인 ProtoBlend를 제안합니다. 먼저, 교사 모델(teacher model)의 지도를 받아 각 원본 비디오에서 높은 신뢰도의 시각적 부분을 선택합니다. 다음으로, 클러스터 기반 프로토타입 할당은 선택된 시각적 부분들을 교사 모델의 특징 공간에서 파티셔닝하고, 클래스 내 클러스터 각각에 대해 하나의 증류 슬롯을 할당합니다. 마지막으로, 각 프로토타입은 해당 클러스터 내의 앵커(anchor)와 블렌딩되고, 그들의 교사 모델 예측값은 동일한 계수를 사용하여 결합하여 혼합된 소스(mixture-source) 기반의 지도 신호를 제공합니다. 네 가지 액션 인식 벤치마크에서 수행한 실험 결과는 ProtoBlend가 증류된 비디오에 대한 반복적인 최적화 없이도 경쟁력 있는 정확도-효율성 균형을 달성한다는 것을 보여줍니다.

Original Abstract

Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most existing approaches synthesize condensed videos through iterative optimization, whose cost is amplified by the temporal dimension. Rather than further reducing the number of optimized variables, we investigate whether effective distilled videos can be constructed without gradient-based optimization of the stored videos. Such a construction-based approach must address three challenges: selecting informative temporal segments, covering diverse intra-class variations under a limited videos-per-class budget, and increasing the information carried by each stored sample. To this end, we propose ProtoBlend, an efficient select-allocate-blend framework. First, teacher-guided temporal clip selection retains a high-confidence segment from each source video. Second, cluster-guided prototype allocation partitions the selected clips in the teacher feature space and assigns one distilled slot to each intra-class cluster. Third, each prototype is blended with an in-cluster anchor, while their teacher predictions are combined using the same coefficient to provide mixture-source supervision. Experiments on four trimmed action-recognition benchmarks demonstrate that ProtoBlend achieves a competitive accuracy-efficiency trade-off without iterative optimization of the distilled videos.

0 Citations
0 Influential
11 Altmetric
55.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!