2608.01672v1 Aug 03, 2026 cs.CL

기억해야 할 것을 학습하기: 문맥 증류를 통한 테스트 시간 학습

Learning What to Remember: Test-Time Training via Context Distillation

Zixin Wen
Zixin Wen
Citations: 121
h-index: 4
Wenhao Chai
Wenhao Chai
Citations: 420
h-index: 8
Hengyu Fu
Hengyu Fu
Citations: 63
h-index: 2
Xingyu Dang
Xingyu Dang
Citations: 102
h-index: 3
Jason D. Lee
Jason D. Lee
Citations: 149
h-index: 2
Ruiming Zhu
Ruiming Zhu
Citations: 171
h-index: 5
Zixuan Wang
Zixuan Wang
Citations: 166
h-index: 6

효과적인 장문맥 모델링은 단순히 과거의 정보를 더 많이 저장하는 것뿐만 아니라, 나중에 유용할 수 있는 정보를 보존하는 것입니다. 테스트 시간 학습(TTT)은 장문맥 모델링을 위한 온라인 파라미터 업데이트를 수행하는 매력적인 접근 방식이지만, 기존의 TTT 방법들은 재구성 또는 온라인 적응 목표만을 최적화하며, 저장된 정보의 미래 유용성을 고려하지 않습니다. 본 연구에서는 테스트 시간 문맥 증류(TTCD)라는 TTT 프레임워크를 제안합니다. TTCD는 제한된 메모리 용량을 미래 사용을 위해 할당하는 자기 지도 학습 목표를 도입합니다. 구체적으로, TTCD는 긴 컨텍스트 창을 가진 '선생님' 모델이 짧은 컨텍스트 창을 가진 '학생' 모델의 가중치를 감독하도록 합니다. 이 두 모델 사이의 은닉 상태 차이는 모델이 미래 토큰 예측에 중요한 문맥 정보를 기억하도록 유도하는 밀집된 자기 지도 신호를 제공합니다. 본 연구에서는 기존 MLP 파라미터를 빠른 가중치로 사용하는 'In-Place TTCD (IP-TTCD)' 변형에 초점을 맞춥니다. 장문맥 언어 모델링 작업에서의 실험 결과, IP-TTCD는 처음부터 사전 학습된 경우 DeltaNet, Gated DeltaNet, 슬라이딩 윈도우 어텐션 및 기존 TTT 방법보다 일관되게 우수한 성능을 보였습니다. 또한, IP-TTCD를 통해 사전 학습된 트랜스포머 모델이 지속적인 사전 학습을 통해 추론 과정에서 파라미터를 조정하여 경량화된 아키텍처 개선만으로 장문맥 기능을 획득할 수 있습니다. 이러한 결과는 TTCD가 아키텍처 기반의 지속적 학습을 향한 중요한 단계임을 보여줍니다.

Original Abstract

Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test-time training (TTT) is an appealing approach that performs online parameter updates for long-context modeling, yet existing TTT methods only optimize either reconstruction or online adaptation objectives without considering the future utility of retained information. In this work, we propose \textbf{T}est-\textbf{T}ime \textbf{C}ontext \textbf{D}istillation (TTCD), a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use. Specifically, TTCD uses a long-window teacher to supervise the fast weights of a short-window student, where the hidden-state discrepancy between them offers a dense, self-supervised signal guiding the model to memorize the contextual information crucial for future token predictions. We focus on an in-place variant: In-Place TTCD (IP-TTCD), which uses the existing MLP parameters as the fast weights. Experiments on long-context language modeling tasks show IP-TTCD consistently outperforms DeltaNet, Gated DeltaNet, sliding-window attention, and TTT when pre-trained from scratch. Furthermore, IP-TTCD allows pre-trained transformer models to adapt their parameters during inference through continual pre-training, gaining long-context capabilities with only a lightweight architectural augmentation. Our results position TTCD as a step toward architectural continual learning.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!