2607.25271v1 Jul 28, 2026 cs.LG

컴퓨팅 및 데이터 최적화 사전 학습의 통합

Bridging Compute- and Data-Optimal Pretraining

K. Hamidieh
K. Hamidieh
Citations: 286
h-index: 7
Tian Qin
Tian Qin
Citations: 95
h-index: 6
David Alvarez-Melis
David Alvarez-Melis
Citations: 133
h-index: 7

기존의 컴퓨팅 최적화 확장 법칙은 무한한 양의 새로운 사전 학습 데이터를 가정하지만, 실제로는 컴퓨팅 자원이 증가하는 속도가 고품질 데이터의 가용성보다 빠른 상황이 점점 더 흔해지고 있습니다. 본 연구에서는 컴퓨팅과 데이터 간의 균형을 맞추는 통합 프레임워크인 Compute-Data (CD) 확장 법칙을 제안합니다. CD 확장 법칙은 데이터를 자유롭게 사용할 수 있는 컴퓨팅 최적화 확장 방식과, 데이터 크기를 고정하고 컴퓨팅 자원을 무한히 늘릴 수 있는 데이터 최적화 방식을 연결합니다. CD 확장 법칙은 토큰 효율성 함수 $η$를 도입하여 기존의 확장 법칙을 확장하며, 이 함수는 새로운 토큰에 비해 파생된 토큰이 가지는 가치를 나타냅니다. 여기에는 다중 에포크 반복 또는 패러프레이징 등을 통해 생성된 토큰이 포함되며, 완벽한 대체재부터 아무런 가치가 없는 경우까지의 범위를 갖습니다. 본 연구에서는 14M에서 600M 파라미터 규모의 모델에 대해 Dolma-3 코퍼스를 사용하여 다중 에포크 반복 및 패러프레이징이라는 두 가지 데이터 확장 전략에 대한 $η$ 값을 측정했습니다. 분석 결과, 토큰 효율성은 일정하지 않으며, 모델 크기, 파라미터당 토큰 비율, 그리고 사용된 파생 데이터의 양에 따라 달라지며, 코퍼스가 확장됨에 따라 포화되는 경향을 보입니다. $η$ 함수의 형태는 모델 크기가 증가하거나 데이터 가용성이 증가함에 따라 컴퓨팅 자원을 데이터를 대체하는 것이 점진적인 효과를 감소시킨다는 것을 의미합니다. 또한, 학습 과정을 컴퓨팅 제한, 데이터 제한, 그리고 모델 제한의 세 가지 운영 모드로 구분하고, 기존의 컴퓨팅 최적화 할당 방식이 실제 환경에서 대부분의 경우 비효율적임을 보여줍니다.

Original Abstract

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!