2606.06479v1 Jun 04, 2026 cs.LG

순환 없이 순환 신경망 사전 학습

Pretraining Recurrent Networks without Recurrence

Akarsh Kumar
Akarsh Kumar
Citations: 64
h-index: 4
Phillip Isola
Phillip Isola
Citations: 479
h-index: 6

순환 신경망(RNN)을 훈련하는 것은 연산의 긴 시퀀스에 걸쳐 신뢰도를 할당해야 하는 문제를 야기합니다. 표준 역전파 알고리즘(BPTT)은 이 문제에 효과적으로 대처하지 못합니다. BPTT는 시간 순서대로 처리되어 병렬성을 제한하며, 기울기 소실 또는 폭주 현상으로 인해 장거리 연관 관계 학습이 어렵습니다. 본 논문에서는 비선형 RNN을 훈련하는 새로운 방법인 Supervised Memory Training (SMT)을 제안합니다. SMT는 RNN의 순환적인 신뢰도 전파를 완전히 회피하고, RNN 훈련을 $(m_t, x_{t+1}) ightarrow m_{t+1}$ 형태의 일단계 메모리 전환 레이블에 대한 지도 학습으로 단순화합니다. SMT는 트랜스포머 기반 인코더를 사용하여 예측 상태 목표를 설정하여 이러한 메모리 레이블을 획득하며, 미래를 예측하는 데 필요한 과거 정보만 유지합니다. SMT는 기억해야 할 내용과 메모리를 업데이트하는 방법을 분리함으로써 RNN 훈련을 시간 병렬 방식으로 수행하고, 모든 토큰 쌍 간에 안정적인 $O(1)$ 길이의 기울기 경로를 제공합니다. 실험 결과, SMT는 언어 모델링 및 픽셀 시퀀스 모델링과 같은 다양한 RNN 아키텍처를 사전 학습할 때 BPTT보다 우수한 성능을 보입니다. SMT는 비선형 RNN이 장거리 의존성을 더 잘 파악하고 병렬로 훈련될 수 있도록 하며, 이는 과거 경험의 시간적 추상화를 구축하는 모델의 확장에 기여할 수 있습니다.

Original Abstract

Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations. Standard backpropagation through time (BPTT) addresses this problem poorly: it is sequential in time, limiting parallelism, and suffers from vanishing or exploding gradients, making long-range associations difficult to learn. We propose Supervised Memory Training (SMT), a method for training nonlinear RNNs that sidesteps recurrent credit propagation entirely by reducing RNN training to supervised learning on one-step memory transition labels $(m_t, x_{t+1}) \rightarrow m_{t+1}$. SMT acquires these memory labels by training a Transformer-based encoder on a predictive state objective--retaining only information from the past necessary to predict the future. By decoupling what to remember from how to update memory, SMT enables time-parallel RNN training with a stable $O(1)$ length gradient path between any two tokens--without ever unrolling the RNN. We find that SMT outperforms BPTT when pretraining various RNN architectures on tasks like language modeling and pixel sequence modeling. SMT enables nonlinear RNNs to better capture long-range dependencies and train in parallel, potentially unlocking the scaling of models that build temporal abstractions of past experience.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!