조건부 ReLU 감소를 이용한 심층 선형 네트워크 이론을 통해 Grokking 현상에서 나타나는 두 가지 학습 단계 분석
Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction
Grokking은 학습 데이터에 적합하는 것과 간단한 기본 규칙을 학습하는 것이 서로 다른 시간 척도에서 발생할 수 있음을 시사합니다. 우리는 이 현상을 분류 손실의 빠른 감소와 학습된 표현의 느린 단순화를 분리함으로써 형식화하고, 이러한 두 가지 정지 시간을 '두 개의 학습 단계'라고 명명했습니다. 심층 선형 네트워크에서, 우리는 '포스트-마진 간격 성장' 또는 '단일 단계 꼬리 수축' 조건이 교차 엔트로피 손실을 로그 시간 척도에서 에플실론 수준으로 감소시키는 것을 보였습니다. 반대로, 계층별 가중치 감소가 존재할 때, 전체 모델의 정규화는 Schatten 유형의 페널티로 표현될 수 있으며, 날카로운 후기 시간 Kurdyka-Lojasiewicz 꼬리를 가지면 이 구조적 에너지는 다항식 시간 척도로 감소합니다. 따라서 두 개의 학습 단계는 데이터 적합과 표현 단순화를 분리합니다. 그런 다음 동일한 메커니즘이 ReLU MLP에서 어떻게 나타나는지 설명합니다. 학습 데이터 세트의 활성화 패턴이 고정된 영역에서는 네트워크가 활성 좌표에서 선형 모델로 축소됩니다. 2계층 ReLU 임베딩 모델에서, 연쇄 법칙 기반 추정은 분류기 헤드가 제어된 하위 스트림 정규화 조건에서 임베딩 블록보다 더 큰 효과적인 그래디언트를 받을 수 있음을 보여줍니다. 이는 분류기가 먼저 적합하고 표현이 나중에 계속 단순화되는 두 단계의 메커니즘을 지지합니다. 우리는 모듈러 덧셈을 주요 실험 설정으로 사용했습니다. 심층 선형 이론은 분석의 엄격한 기반을 제공합니다. 그러나 ReLU 결과는 경험적 행동을 설명하는 조건부 감소 형태로 제기되며, 비선형 학습 동역학에 대한 전역적인 증명을 주장하지 않습니다.
Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomenon by separating the fast decay of the classification loss from the slower simplification of the learned representation, and we call the resulting pair of stopping times two training clocks. For deep linear networks, we show that a post-margin gap-growth or one-step tail-contraction condition reduces the cross-entropy loss to level epsilon on a logarithmic time scale. In contrast, when layerwise weight decay is present, the induced regularization on the end-to-end map can be expressed as a Schatten-type penalty; under a sharp late-time Kurdyka-Lojasiewicz tail, this structural energy closes on a polynomial time scale. The two clocks, therefore, separate fitting from representation simplification. We then explain how the same mechanism can appear in ReLU MLPs. In regions where the activation patterns on the training set remain fixed, the network reduces to a linear model in the active coordinates. In a two-layer ReLU embedding model, chain-rule estimates further show that the classifier head can receive larger effective gradients than the embedding block under controlled downstream norms. This supports a two-stage mechanism in which the classifier fits first, while the representation continues to simplify later. We use modular addition as the main experimental setting. The deep linear theory provides the rigorous core of the analysis. But the ReLU results are formulated as conditional reductions that account for empirical behavior without claiming a global proof for nonlinear training dynamics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.