학습의 두 가지 속도: Grokking과 Double Descent 현상의 표현-읽기 분해
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent
심층 신경망 학습 과정에서 일반화 성능을 모니터링하기 위해 일반적으로 손실(loss)과 정확도를 사용합니다. 그러나 grokking과 epoch-wise double descent이라는 두 가지 잘 알려진 현상은 이러한 관찰에 복잡성을 더합니다. grokking에서는 학습 손실이 빠르게 감소하는 반면, 테스트 성능은 상당한 지연 후에 갑자기 향상됩니다. epoch-wise double descent에서는 학습 손실이 단조롭게 감소하는 동안 테스트 손실 또는 오류는 증가했다가 감소하는 경향을 보입니다. 기존 연구들은 종종 특정 작업에 국한되어 있으며, 현실적인 작업과 아키텍처 전반에 걸쳐 이러한 현상을 진단하고 설명할 수 있는 일반적인 분석 프레임워크가 부족합니다. 우리는 이 문제를 해결하기 위해 학습 역학의 기반이 되는 두 가지 경쟁적 과정을 분석합니다. 즉, 인코더에서의 표현 학습과 최종 분류기에서의 읽기(readout) 보정입니다. 표현 기하학, 신경 접선 커널 및 선형 프로빙 도구를 사용하여 학습 과정 전반에 걸쳐 이러한 두 가지 프로세스가 모두 활성화되어 있으며, 상대적인 속도의 변화가 예상치 못한 일반화 역학을 초래한다는 것을 보여줍니다. 우리는 다양한 작업과 아키텍처에서 grokking 현상에 representation-readout 분해를 적용하여 읽기가 grokking 시작 전에 학습 데이터에 편향되어 있으며, 표현 학습은 점진적이지만 존재하지 않는다는 'lazy-to-rich' 설명과는 달리 나타난다는 것을 발견했습니다. 또한 이 프레임워크는 가짜 일반화와 진정한 일반화를 구별하는 지표를 제공합니다. 이전 보고된 MNIST grokking 예제 및 epoch-wise double descent 예제에서, 겉보기에는 지연되거나 비단조적인 일반화 현상은 표준이 아닌 학습 방법을 통해 유발되는 표현의 저하 및 읽기 불일치로 인해 발생하는 것으로 나타났습니다. 이러한 결과들을 종합하면 representation-readout 분해는 학습 역학을 이해하고 해석 가능성 연구를 위한 기본 알고리즘을 밝히는 데 유용한 상위 수준 프레임워크임을 입증합니다.
Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improves abruptly only after a long delay; in epoch-wise double descent, train loss decreases monotonically while test loss or error rises and falls. Existing accounts are often task-specific, and a task-agnostic analysis framework for diagnosing and explaining these phenomena across realistic tasks and architectures is missing. We address this challenge by analyzing two competing processes that underlie learning dynamics: representation learning in the encoder and readout calibration in the final classifier. Using tools from representational geometry, neural tangent kernels, and linear probing, we show that both processes are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent, contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalization is shown to arise from representation degradation and readout misalignment induced by non-standard training recipes. Together, these results establish the representation-readout decomposition as a top-down framework for understanding learning dynamics and revealing underlying algorithms for interpretability research.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.