컴퓨팅 및 데이터 최적화 사전 학습의 통합
Bridging Compute- and Data-Optimal Pretraining
기존의 컴퓨팅 최적화 확장 법칙은 무한한 양의 새로운 사전 학습 데이터를 가정하지만, 실제로는 컴퓨팅 자원이 증가하는 속도가 고품질 데이터의 가용성보다 빠른 상황이 점점 더 흔해지고 있습니다. 본 연구에서는 컴퓨팅과 데이터 간의 균형을 맞추는 통합 프레임워크인 Compute-Data (CD) 확장 법칙을 제안합니다. CD 확장 법칙은 데이터를 자유롭게 사용할 수 있는 컴퓨팅 최적화 확장 방식과, 데이터 크기를 고정하고 컴퓨팅 자원을 무한히 늘릴 수 있는 데이터 최적화 방식을 연결합니다. CD 확장 법칙은 토큰 효율성 함수 $η$를 도입하여 기존의 확장 법칙을 확장하며, 이 함수는 새로운 토큰에 비해 파생된 토큰이 가지는 가치를 나타냅니다. 여기에는 다중 에포크 반복 또는 패러프레이징 등을 통해 생성된 토큰이 포함되며, 완벽한 대체재부터 아무런 가치가 없는 경우까지의 범위를 갖습니다. 본 연구에서는 14M에서 600M 파라미터 규모의 모델에 대해 Dolma-3 코퍼스를 사용하여 다중 에포크 반복 및 패러프레이징이라는 두 가지 데이터 확장 전략에 대한 $η$ 값을 측정했습니다. 분석 결과, 토큰 효율성은 일정하지 않으며, 모델 크기, 파라미터당 토큰 비율, 그리고 사용된 파생 데이터의 양에 따라 달라지며, 코퍼스가 확장됨에 따라 포화되는 경향을 보입니다. $η$ 함수의 형태는 모델 크기가 증가하거나 데이터 가용성이 증가함에 따라 컴퓨팅 자원을 데이터를 대체하는 것이 점진적인 효과를 감소시킨다는 것을 의미합니다. 또한, 학습 과정을 컴퓨팅 제한, 데이터 제한, 그리고 모델 제한의 세 가지 운영 모드로 구분하고, 기존의 컴퓨팅 최적화 할당 방식이 실제 환경에서 대부분의 경우 비효율적임을 보여줍니다.
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.