2608.01488v1 Aug 02, 2026 cs.CV

컴팩트하고 통합적인 다중 모드 추적을 향하여: 지식 증류와 구조적 가지치기를 결합

Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

Chuanguang Yang
Chuanguang Yang
Citations: 359
h-index: 11
Yingli Tian
Yingli Tian
Citations: 197
h-index: 7
Shiping Wen
Shiping Wen
Citations: 27
h-index: 2
Tingwen Huang
Tingwen Huang
Citations: 53
h-index: 4
Huiran Duan
Huiran Duan
Citations: 4
h-index: 1
Zhulin An
Zhulin An
Citations: 41
h-index: 4
Yuqi Li
Yuqi Li
Citations: 194
h-index: 8
Yuedong Tan
Yuedong Tan
Citations: 133
h-index: 5
Weilun Feng
Weilun Feng
Citations: 132
h-index: 7
Zongwei Wu
Zongwei Wu
Citations: 873
h-index: 17

통합 다중 모드 객체 추적은 RGB, 열화상, 깊이 등 상호 보완적인 센서 데이터를 활용하여 뛰어난 안정성을 달성했지만, 최첨단 모델의 높은 계산 비용으로 인해 자원 제약 환경의 엣지 장치에 적용하기 어렵습니다. 본 연구에서는 예측 헤드를 중요한 효율성 병목 지점으로 밝혀내고, 디코더 아키텍처를 전략적으로 간소화하여 실시간 추론 가능성을 높이는 동시에 경량 모델(학생)과 고성능 모델(선생) 간의 성능 격차를 야기합니다. 이 문제를 해결하기 위해 17가지 증류 방법을 체계적으로 분석하고, Dual-Alignment Distillation 프레임워크를 제안합니다. 핵심 아이디어는 효과적인 압축을 위해서는 지식 전달을 두 가지 상호 보완적인 흐름으로 분리해야 한다는 것입니다. (1) Spatial Representation Alignment: 특징 증류를 사용하여 학생 모델이 전경 객체에 대한 공간적 집중력을 향상시키도록 합니다 (“어디를 추적할 것인가”). (2) Semantic Distribution Alignment: 로짓 기반 증류를 활용하여 의사 결정 경계를 조정하고, 판별력 있는 암묵적인 지식을 전달합니다 (“무엇을 추적할 것인가”). 다섯 가지 벤치마크에서 실시한 광범위한 실험 결과, 제안하는 방법이 복잡한 최첨단 방법보다 훨씬 뛰어난 성능을 보였습니다. 특히, 증류된 모델은 RGBT234 데이터셋에서 91.5%의 MPR(Mean Precision)을 달성하고, 단일 RTX 4090 GPU에서 54 FPS로 작동하며, 이는 선생님 모델 대비 5배 빠른 속도이며, 우수한 정확도를 유지합니다.

Original Abstract

Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy computational burden of state-of-the-art models hinders their deployment on resource-constrained edge devices. In this work, we identify the prediction head as a critical but often overlooked efficiency bottleneck. By strategically streamlining the decoder architecture, we unlock the potential for real-time inference but simultaneously introduce a capacity gap between the lightweight student and the heavy teacher. To resolve this, we conduct a systematic analysis of 17 distillation strategies and introduce a Dual-Alignment Distillation framework. Our key insight is that effective compression requires decoupling knowledge transfer into two complementary streams: (1) Spatial Representation Alignment, which employs feature distillation to sharpen the student's spatial focus on foreground targets ("Where to track"); and (2) Semantic Distribution Alignment, which utilizes logit-based distillation to align decision boundaries and transfer discriminative dark knowledge ("What to track"). Extensive experiments across five benchmarks demonstrate that our approach significantly outperforms complex state-of-the-art methods. Notably, our distilled model achieves 91.5% MPR on RGBT234 and operates at 54 FPS on a single RTX 4090, representing a 5x speedup over the teacher model while maintaining superior accuracy.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!