2606.16112v1 Jun 15, 2026 cs.LG

Norm-agnostic 잔차 네트워크를 이용한 적응적 깊이 확장

Scaling Adaptive Depth with Norm-Agnostic Residual Networks

Beren Millidge
Beren Millidge
Citations: 2,358
h-index: 26
Tomas Figliolia
Tomas Figliolia
Citations: 16
h-index: 2

잔차 구조는 딥러닝에서 널리 사용되지만, 미묘한 구조적인 한계점을 가지고 있습니다. 즉, 잔차 스트림의 노름(norm) 값이 깊이에 따라 급격하게 증가하는 경향이 있습니다. 그 결과, 후반 레이어에서의 업데이트는 누적된 잔차 상태에 비해 상대적으로 작아져 표현 학습에 미치는 영향이 줄어들고, 모델의 깊이를 확장했을 때 얻을 수 있는 이점을 제한합니다. 이러한 문제를 해결하기 위해, 우리는 노름(norm) 값에 영향을 받지 않는 잔차 구조인 NAG를 제안합니다. NAG는 잔차 스트림에서 크기 정보와 방향 정보를 분리하여, 전체적인 깊이에 걸쳐 각 레이어가 갖는 의미있는 기여도를 유지하고, 잔차 노름의 증가로 인해 후반 레이어의 업데이트가 체계적으로 억제되는 현상을 방지합니다. 중요한 점은 NAG는 매우 적은 수의 추가 파라미터를 사용하며, 간단한 연산을 통해 효율적인 학습이 가능하도록 설계되었다는 것입니다. 실험 결과, NAG 기반 구조는 기존 Transformer 모델보다 우수한 성능을 보였으며, 특히 깊이가 증가할수록 성능 향상이 두드러졌습니다. 또한, 노름(norm)에 영향을 받지 않는 특성 덕분에, 어텐션(attention) 및 MLP 레이어를 적응적으로 건너뛰는 interpretable한 Mixture-of-Depths (MoD) 메커니즘을 구현할 수 있었습니다. MoD 메커니즘은 단순히 후처리 과정에서 정확도와 계산량 사이의 균형을 맞추는 데 사용될 뿐만 아니라, 사전 학습 단계에서의 모델 확장 전략으로 활용될 수 있습니다. 즉, 고정된 FLOP(floating-point operations) 환경에서 각 토큰에 대한 연산 비용을 줄여 확보한 계산 자원을 더 많은 토큰을 사용하여 학습하는 데 재투자함으로써, 전체 파라미터 수와 KV 캐시 예산을 유지하면서 더 깊은 모델을 학습할 수 있습니다. 실험 결과, 약 20%-25% 정도의 MoD 비율을 적용했을 때, 동일한 학습 계산량을 사용하면서 기존의 full-depth baseline 모델과 유사한 성능을 보이면서도, 실제로 실행되는 레이어 파라미터 및 forward-pass FLOPs 수를 크게 줄일 수 있었습니다. 이러한 결과는 깊이에서의 희소성(sparsity)이 고정된 계산량 환경에서 모델 학습의 새로운 확장 축(scaling axis)이 될 수 있음을 보여주며, 매우 깊으면서도 FLOP 효율적인 모델을 개발할 수 있는 가능성을 제시합니다.

Original Abstract

Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual stream can grow rapidly with depth. As a result, updates from later layers become small relative to the accumulated residual state. This reduces their impact on the representation and limits the benefits of scaling models in depth. To address this, we introduce NAG, a norm-agnostic residual architecture that separates magnitude from directional information in the residual stream, preserving meaningful layer contributions throughout depth and preventing later updates from being systematically suppressed by residual-norm growth. Importantly, NAG introduces only a negligible number of additional parameters and relies on simple operations that are easily kernel-fusible, preserving training efficiency in practice. We show that this architecture outperforms baseline Transformers, with gains that increase substantially as depth grows, enabling effective training of much deeper models. The norm-agnostic formulation also leads to an interpretable Mixture-of-Depths (MoD) mechanism that adaptively skips both attention and MLP layers. Beyond serving as a post-training accuracy-compute tradeoff, this mechanism can be used as a pretraining-time scaling strategy: under iso-FLOP training, compute saved by reducing per-token forward-pass cost can be reinvested into training on more tokens while keeping the total parameter count and KV-cache budget fixed. In our experiments, moderate Mixture-of-Depths rates of approximately 20%-25% match full-depth baseline performance under equal training compute while substantially reducing the number of executed layer parameters and forward-pass FLOPs. These results identify sparsity in depth as a new scaling axis for fixed-compute training, enabling very deep yet FLOP-efficient models.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!