2606.19179v1 Jun 17, 2026 cs.LG

확률적 모멘텀 방법의 계산 효율성과 순차 실행 시간 간의 상호 관계 연구

Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods

Alexandru Meterez
Alexandru Meterez
Harvard University
Citations: 192
h-index: 6
Depen Morwani
Depen Morwani
Citations: 578
h-index: 9
Sham M. Kakade
Sham M. Kakade
Citations: 5
h-index: 1
P. Nair
P. Nair
Citations: 38
h-index: 4

헤비 볼(HB), 네스테로프 모멘텀, 그리고 가속화된 확률적 경사 하강법(ASGD)의 변형과 같은 확률적 모멘텀 방법들은 현대적인 훈련 과정에서 널리 사용되지만, 이들의 장점은 순차 실행 시간 (목표 정확도에 도달하는 데 필요한 반복 횟수)과 계산 효율성 (총 그래디언트 쿼리 또는 FLOP 비용의 역수)이라는 두 가지 중요한 요소에 따라 달라집니다. 큰 배치 크기는 계산 효율성을 저하시키지 않으면서 순차 실행 시간을 줄이지만, 이는 수축 간격이 배치 크기에 대해 선형적으로 증가할 때만 가능합니다. 본 연구에서는 가우시안 공변량을 갖는 일관된 선형 회귀 문제에 대한 확률적 HB와 ASGD를 분석하고, 배치 크기 변화에 따른 하한 경계를 유도했습니다. 첫 번째 결과로, HB는 임의의 스펙트럼에서 SGD보다 계산 효율성을 향상시키지 못하며, 오히려 더 넓은 범위의 배치 크기에 대해 SGD 수준의 계산 효율성을 유지합니다. 이를 통해 더 큰 배치 크기를 사용하여 순차 실행 시간을 줄일 수 있으며, HB가 결정론적인 가속화 수준에 도달할 때까지 이러한 효과가 지속됩니다. 이 범위는 SGD의 임계 배치 크기보다 $\sqrt{\kappa}$배 더 클 수 있습니다. ASGD의 경우, 그 효과는 스펙트럼 특성에 더욱 의존적입니다. 급격하게 감소하는 거듭제곱 법칙 스펙트럼에서, ASGD는 작은 배치 크기에서 HB/SGD보다 계산 효율성이 높지만, 배치 크기가 증가함에 따라 이러한 계산 효율성 이점을 순차 실행 시간 개선으로 대체합니다. 합성 선형 회귀 실험을 통해 이러한 경향을 확인했으며, 이는 천천히 감소하는 스펙트럼에서는 ASGD와 HB가 거의 유사한 성능을 보이고, 급격하게 감소하는 스펙트럼에서는 예측된 계산 효율성-순차 실행 시간 간의 상호 관계를 나타냅니다.

Original Abstract

Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are widely used in modern training, but their stochastic benefits depend on two distinct quantities: serial runtime, the number of iterations needed to reach a target accuracy, and compute efficiency (CE), the inverse total gradient-query or FLOP cost. Larger batches reduce serial runtime without hurting CE only when the contraction gap grows linearly with batch size. We study stochastic HB and ASGD for consistent linear regression with Gaussian covariates and prove finite-dimensional, discrete-time lower bounds on their batch-size tradeoffs. Our first result shows that HB does not improve the CE frontier over SGD for arbitrary spectra; rather, it preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale. This window can be a factor $\sqrtκ$ larger than the SGD critical batch size. For ASGD, the picture is more spectrum-dependent: for rapidly decaying power-law spectra, ASGD improves small-batch CE over HB/SGD, but as batch size grows it trades this CE advantage for improved serial runtime. Synthetic linear-regression experiments verify these qualitative regimes, including near-overlap of ASGD and HB for slowly decaying spectra and the predicted CE--serial tradeoff for rapidly decaying spectra.

1 Citations
0 Influential
4.5 Altmetric
23.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!