2608.13156v1 Aug 13, 2026 cs.AI

LLM의 정규화 배치에 대한 재고: 커리큘럼 기반 심층 모델에서 후정규화(Post-Norm) 연구

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

Naiqiang Tan
Naiqiang Tan
Citations: 377
h-index: 6
Jun Fang
Jun Fang
Citations: 120
h-index: 5
Yadong Wang
Yadong Wang
Citations: 2
h-index: 1
Xiang Chen
Xiang Chen
Citations: 2
h-index: 1
Rui Liu
Rui Liu
Citations: 86
h-index: 4
Sheng Ren
Sheng Ren
Citations: 10
h-index: 2
Jiangang Kong
Jiangang Kong
Citations: 47
h-index: 2
Jun Wang
Jun Wang
Citations: 24
h-index: 2
Kai Chen
Kai Chen
Citations: 26
h-index: 2
Li Liang
Li Liang
Citations: 0
h-index: 0

현대 트랜스포머 모델에서 사전정규화(Pre-norm)는 전체 깊이 모델의 공동 최적화를 용이하게 하기 때문에 표준적인 정규화 배치 방식입니다. 본 논문에서는 심층 구조가 커리큘럼을 통해 도입될 때 이러한 선호도가 유지되는지 질문합니다. 커리큘럼 기반 심층 성장에서, 각 추가된 블록은 훈련된 접두부(prefix)에 의해 생성된 경계 표현을 입력받으므로, 정규화 배치는 순방향 조건 설정과 관련됩니다. 따라서 본 연구에서는 배치 방식과 훈련 커리큘럼이 상호 작용하는지 테스트합니다. Qwen3-8B 모델을 교사 모델로 사용하고 9개의 레이어를 가진 학생 모델을 사용하여 통제된 지식 증류(distillation) 실험을 진행한 결과, 공동 훈련 하에서는 사전정규화와 후정규화가 $0.0004$의 미미한 차이만 보였습니다. 그러나 커리큘럼 성장을 사용하는 경우, 후정규화는 사전정규화보다 $0.0328$만큼 성능이 향상되었으며, 이는 매우 큰 격차입니다. 학생 모델의 활성 레이어 토큰을 사용하여 통제된 실험에서도 후정규화가 여전히 더 우수했으며, 이는 계산량만으로는 설명할 수 없는 결과입니다. 훈련 과정 중 정규화 방식의 순위가 바뀌는 것을 관찰했는데, 블록이 추가되기 시작하면 후정규화가 앞서게 됩니다. 단일 블록 사용 및 동결(freeze) 실험을 통해 순위 변화가 블록 추가 자체에 국한되어 있으며, 얕은 블록의 품질이나 재훈련과는 관련이 없음을 확인했습니다. 경계 분석 결과, 후정규화는 안정적인 잔차 스케일을 갖는 반면, 사전정규화는 구조적 토큰 스케일 드리프트(drift)를 나타냅니다. 또한 고정된 배치 크기에서 최종 커리큘럼 성장 블록은 거의 동일 변환(identity mapping)에 가깝습니다. 이러한 관찰 결과들은 새로운 블록이 추가될 때 경계 스케일에 대한 조건 설정과 함께, 단계별 순위 변화를 뒷받침합니다. 본 연구의 결과는 이 지식 증류 환경에서 정규화 배치 방식과 훈련 커리큘럼을 상호 연관된 설계 선택으로 고려해야 함을 시사합니다.

Original Abstract

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!