2604.25779v1 Apr 28, 2026 cs.LG

지속적인 기울기 정렬이 다단계 환경에서 잠재적 학습을 매개: MNIST 보조 로짓 증류 실험을 통한 증거

Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment

Shivam Arora
Shivam Arora
Citations: 11
h-index: 2
Chayanon Kitkana
Chayanon Kitkana
Citations: 3
h-index: 1

MNIST 보조 로짓 증류 실험에서, 학생 모델은 특정 클래스에 속하지 않는 로짓 데이터만을 사용하여 학습함에도 불구하고, 잠재적 학습(subliminal learning)이라는 현상을 통해 의도하지 않은 교사 모델의 특성을 습득할 수 있습니다. 단일 단계 기울기 하강(gradient descent)을 가정할 때, 잠재적 학습 이론은 이러한 효과를 특성과 증류 기울기 간의 정렬(alignment)로 설명하지만, 다단계 환경에서는 이러한 정렬이 유지되는지 보장하지 않습니다. 본 연구에서는 실험적으로 기울기 정렬이 훈련 과정 전반에 걸쳐 약하지만 일관되게 긍정적인 값을 유지하며, 특성 획득에 인과적으로 기여한다는 것을 입증했습니다. 또한, 정렬을 약화시키는 완화 방법인 '림날(liminal) 훈련'이 작동하는 원리가 정렬 감소에 있음을 확인했으며, 이 방법은 본 실험 환경에서 특성 획득을 완전히 막지 못한다는 것을 보여주었습니다. 이러한 결과는 1차 기울기(first-order drive)가 우세한 환경에서, 이러한 방식으로 작동하는 완화 방법들이 특성 획득을 안정적으로 억제하지 못할 수 있음을 시사합니다.

Original Abstract

In the MNIST auxiliary logit distillation experiment, a student can acquire an unintended teacher trait despite distilling only on no-class logits through a phenomenon called subliminal learning. Under a single-step gradient descent assumption, subliminal learning theory attributes this effect to alignment between the trait and distillation gradients, but does not guarantee that this alignment persists in a multi-step setting. We empirically show that gradient alignment remains weakly but consistently positive throughout training and causally contributes to trait acquisition. We show that a mitigation method called liminal training works by attenuating the alignment and fails to stop trait acquisition in this setup. These results suggest that mitigation methods that operate in this regime may not reliably suppress trait acquisition when the first-order drive dominates.

2 Citations
0 Influential
1 Altmetric
7.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!