2606.19781v1 Jun 18, 2026 hep-ex

사전 학습 데이터 구성에 따른 모델 확장 법칙 연구

Towards Engineering Scaling Laws with Pretraining Data Composition

Jan-Lucas Uslu
Jan-Lucas Uslu
Citations: 224
h-index: 4
K. Greif
K. Greif
Citations: 246
h-index: 10
D. Whiteson
D. Whiteson
Citations: 49,740
h-index: 95
B. Nachman
B. Nachman
Citations: 27,485
h-index: 82

신경망 확장 법칙은 모델 성능이 컴퓨팅 자원, 모델 크기 및 데이터 세트 크기에 따라 지수적으로 향상되는 현상을 설명합니다. 이 관계는 대규모 언어 모델에서 잘 확립되었지만, 최근에는 입자 물리학 분야의 대규모 모델에서도 나타나고 있습니다. 언어와 마찬가지로, 실증 연구 결과에 따르면 성능은 지수 함수 형태로 확장됩니다. 그러나 자연어 또는 이미지 도메인과는 달리, 입자 물리학에서는 고정밀 시뮬레이터가 저렴한 비용으로 합성 데이터를 생성할 수 있습니다. 이러한 특성은 추가 파라미터보다 추가 데이터가 더 저렴한 확장 단계를 선호하며, 사전 학습 데이터 세트 자체를 설계하여 확장에 영향을 미칠 수 있도록 합니다. 본 연구에서는 고에너지 입자빔 충돌 시 발생하는 강입자 제트를 분류하는 작업에서, 사전 학습 데이터의 다양성과 하위 분류 작업과의 정렬성을 높임으로써, 더 큰 모델보다는 더 많은 데이터를 필요로 하는 확장 방식으로 성능을 향상시킬 수 있음을 보여줍니다.

Original Abstract

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established for large language models, these relationships are emerging for large models in particle physics. As with language, empirical studies show that the performance scales as a power law. However, unlike natural language or image domains, fundamental physics has high-fidelity simulators that produce synthetic data cheaply. This favors scaling regimes where additional data is cheaper than additional parameters, and allows the pretraining dataset itself to be engineered to influence the scaling. For the task of classifying hadronic jets produced in collisions of high-energy particle beams, we show that the scaling behavior can be engineered towards requiring more data rather than larger models by inclusion of pretraining data which is more diverse and better aligned with the downstream classification task.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!