2602.10545v1 Feb 11, 2026 cs.LG

마이크로초(μs) 단위의 스케일링을 이용한 소형 모델: 원칙에 기반한 초기화 및 하이퍼파라미터 전이

$μ$pscaling small models: Principled warm starts and hyperparameter transfer

Yuxin Ma
Yuxin Ma
Citations: 7
h-index: 1
Nan Chen
Nan Chen
Citations: 70
h-index: 4
M. D'iaz
M. D'iaz
Citations: 73
h-index: 5
Soufiane Hayou
Soufiane Hayou
Citations: 12
h-index: 2
Dmitriy Kunisky
Dmitriy Kunisky
Citations: 596
h-index: 10
Soledad Villar
Soledad Villar
Citations: 18
h-index: 2

최근의 대규모 신경망은 다양한 추론 비용에 맞춰 여러 크기로 훈련 및 출시되는 경우가 많습니다. 효율성을 높이기 위해 최근 연구에서는 모델 업스케일링(model upscaling)이 탐구되었습니다. 이는 더 큰 모델을 훈련된 더 작은 모델로부터 초기화하여 지식을 이전하고 수렴 속도를 가속화하는 방법입니다. 그러나 이 방법은 목표 업스케일링 모델 크기에 맞춰 조정해야 하는 하이퍼파라미터에 민감하며, 직접적으로 이를 조정하는 것은 비용이 매우 많이 듭니다. 현재 가장 흔히 사용되는 방법인, 더 작은 모델에서 하이퍼파라미터를 조정하고 하이퍼파라미터 스케일링 법칙을 통해 결과를 추정하는 방식이 업스케일링을 사용할 때도 여전히 유효한지 명확하지 않습니다. 본 연구에서는 모델의 폭(width)에 대한 업스케일링에 대한 원칙적인 접근 방식을 제시하고, 이 환경에서 하이퍼파라미터를 효율적으로 조정하는 방법을 다룹니다. 첫째, μP 아키텍처 및 임의 차원 아키텍처에서 영감을 받아, 다양한 아키텍처 및 최적화 기법에 적용 가능한 일반적인 업스케일링 방법을 제안합니다. 이 방법은 모델이 확장된 버전과 동일하다는 것을 보장하는 이론적 근거를 가지고 있으며, 무한 폭(infinite-width)의 극한에 대한 엄격한 분석을 가능하게 합니다. 둘째, μTransfer 이론을 확장하여 제안하는 방법으로 업스케일링된 모델을 위한 하이퍼파라미터 전이 기법을 개발하고, 실제 데이터셋 및 아키텍처에서 이 방법이 효과적임을 경험적으로 입증합니다.

Original Abstract

Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent work has explored model upscaling: initializing larger models from trained smaller ones in order to transfer knowledge and accelerate convergence. However, this method can be sensitive to hyperparameters that need to be tuned at the target upscaled model size, which is prohibitively costly to do directly. It remains unclear whether the most common workaround -- tuning on smaller models and extrapolating via hyperparameter scaling laws -- is still sound when using upscaling. We address this with principled approaches to upscaling with respect to model widths and efficiently tuning hyperparameters in this setting. First, motivated by $μ$P and any-dimensional architectures, we introduce a general upscaling method applicable to a broad range of architectures and optimizers, backed by theory guaranteeing that models are equivalent to their widened versions and allowing for rigorous analysis of infinite-width limits. Second, we extend the theory of $μ$Transfer to a hyperparameter transfer technique for models upscaled using our method and empirically demonstrate that this method is effective on realistic datasets and architectures.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!