DeepLoop: 루프 트랜스포머를 위한 깊이 스케일링
DeepLoop: Depth Scaling for Looped Transformers
루프 트랜스포머는 물리적 블록의 작은 스택을 여러 번 반복하여 순차적인 계산을 확장하며, 저장되는 파라미터 수를 늘리지 않고 언롤된 깊이를 증가시킵니다. 이러한 재사용은 잔차 스케일링 문제를 변화시키는데, 일반적인 트랜스포머에서는 각 잔차 분기마다 자체 파라미터 업데이트를 받지만, 루프 트랜스포머에서는 하나의 공유 업데이트가 반복적인 방문으로부터 발생하는 기울기를 집계하고, 다음 선형 순방향 패스에서 동일한 방문에 의해 다시 사용됩니다. 우리는 이러한 결합된 깊이 효과를 방문 정렬 계수 $κ_R$에 의해 제어되는 1차 퍼터베이션 경계를 통해 공식화합니다. 이 경계는 방문이 상관 관계가 없을 때 DeepNorm 지수를 회복하지만, 보수적인 정렬 상태에서는 물리적 깊이가 고정된 상태에서 루프 횟수가 증가함에 따라 지수가 $1/4$에서 $1/2$로 증가해야 합니다. 결과적으로, extbf{DeepLoop} 방법은 Post-LN DeepNorm 아키텍처를 유지하고 언롤된 깊이 $N$에 대해 $α=(2N)^{1/2}$ 및 $β=(8N)^{-1/2}$로 설정합니다. GPT-2 small 및 GPT-2 medium 규모의 GPT 스타일 루프 언어 모델에서, DeepLoop는 물리적 블록이 재방문되지 않는 경우 성능 변화가 없으며, 순환 깊이가 활성화되면 검증 손실과 다운스트림 정확도를 향상시킵니다. 이러한 결과는 안정적인 순환 깊이를 위해서는 파라미터 방문을 고려하는 잔차 스케일링 규칙이 필요하며, 단순히 레이어 수를 기준으로 할 수 없음을 보여줍니다.
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient $κ_R$. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from $1/4$ to $1/2$ as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets $α=(2N)^{1/2}$ and $β=(8N)^{-1/2}$ for unrolled depth $N$. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.