2407.13911v5 Jul 18, 2024 cs.CV

분리된 프롬프트 기반의 리허설 없는 클래스 증분 학습을 위한 지속적인 지식 전달 학습

Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting

Yu Xiang
Yu Xiang
Citations: 44
h-index: 3
Qifan Zhang
Qifan Zhang
Citations: 63
h-index: 3
Yunhui Guo
Yunhui Guo
Citations: 44
h-index: 2

프롬프트 기반의 지속적 학습은 사전 훈련된 비전 트랜스포머(ViT) 백본을 고정하고 학습 가능한 프롬프트를 조정함으로써, 리허설이 필요 없는 클래스 증분 학습에서 뛰어난 성능을 보여주었습니다. 그러나 백본 모델의 크기가 성능에 미치는 영향은 아직 충분히 연구되지 않았습니다. 본 논문에서는 더 큰 ViT 백본이 지속적인 학습 성능 측면에서 일관되게 더 나은 결과를 가져온다는 것을 확인하고, 이러한 능력을 더 작은 모델로 이전하는 방법을 연구합니다. 본 논문에서는 리허설이 없는 프롬프트 기반의 지속적 학습에 대한 새로운 지식 전달 방법인 지속적인 지식 전달 학습(Continual Distillation Learning, CDL)을 제안합니다. 기존의 지식 전달 방법은 CDL에서 제한적인 효과만 가져다줍니다. 그 이유는 작업별 프롬프트가 지속적인 적응과 지식 전달이라는 두 가지 역할을 동시에 수행해야 하지만, 작업 간 지식 이전 메커니즘이 부족하기 때문입니다. 이러한 문제를 해결하기 위해, 본 논문에서는 분리된 지식 전달 학습(Decoupled Continual Distillation Learning, D-CDL)을 제안합니다. D-CDL은 지속적인 지식 전달 프롬프트(KD-prompts)와 전용 KD 브랜치를 도입하여 지식 전달과 작업 적응을 명시적으로 분리합니다. 제안된 KD-프롬프트는 여러 작업에 걸쳐 교사 모델의 지식을 운반하는 역할을 하며, 원래의 프롬프트는 지속적인 학습을 담당합니다. D-CDL은 간단하고 일반적이며 다양한 프롬프트 기반 지속적 학습 프레임워크에 통합될 수 있습니다. Split CIFAR-100 및 Split ImageNet-R 데이터셋에서 네 가지 대표적인 지속적 학습 방법에 대한 광범위한 실험 결과, D-CDL은 기존의 지식 전달 방법보다 우수한 성능을 보이며 다양한 교사-학생 설정에서 학생 모델의 성능을 크게 향상시킵니다.

Original Abstract

Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transformer (ViT) backbone. However, the effect of backbone scale remains underexplored. We observe that larger ViT backbones consistently yield better continual learning performance, which motivates us to study how to transfer such capability from a larger model to a smaller one. In this paper, we introduce Continual Distillation Learning (CDL), a new setting for knowledge distillation in rehearsal-free prompt-based continual learning. We show that conventional distillation methods provide only limited gains in CDL, mainly because task-specific prompts are forced to encode both continual adaptation and distillation knowledge, while lacking a persistent mechanism for cross-task knowledge transfer. To address this problem, we propose Decoupled Continual Distillation Learning (D-CDL), which introduces persistent Knowledge-Distillation prompts (KD-prompts) and a dedicated KD branch to explicitly decouple distillation from task adaptation. The proposed KD-prompts are propagated across tasks as a global carrier of teacher knowledge, while the original prompts remain responsible for continual learning. D-CDL is simple, general, and can be integrated into various prompt-based continual learning frameworks. Extensive experiments on Split CIFAR-100 and Split ImageNet-R across four representative continual learning methods show that D-CDL consistently outperforms existing distillation baselines and substantially improves student performance under different teacher-student settings.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!