2608.03197v1 Aug 04, 2026 cs.LG

샤프니스 인식 최소화의 암묵적 평탄성 편향에 대한 연구: 정량적인 하이퍼파라미터 경계를 갖는 선형 안정성 분석

On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

Junbiao Pang
Junbiao Pang
Citations: 18
h-index: 2
Jiaxin Deng
Jiaxin Deng
Citations: 33
h-index: 3

샤프니스 인식 최소화(SAM)는 로컬 적대적 변화에 강건한 파라미터를 탐색하여 일반화 성능을 향상시키지만, SAM이 평탄한 최소값을 지향하는 암묵적인 메커니즘은 아직 명확하지 않습니다. 특히, 섭동 반경 $ρ$는 일반적으로 SAM이 샤프니스를 측정하는 이웃을 정의함에도 불구하고 독립적인 조정 파라미터로 취급됩니다. 본 논문에서는 미니 배치 SAM을 사용하여 보간 최소점 근처에서 선형 안정성을 분석합니다. 로컬 선형화 및 기울기 노이즈 정렬 가정을 바탕으로, 모든 선형적으로 안정적인 최소점이 $λ_{ ext{max}} ≤ oot 3 elax bΓ/(2ρη^2)$를 만족함을 증명합니다. 여기서 $λ_{ ext{max}}$는 최대 헤세 행렬 고유값, $b$는 배치 크기, $η$는 학습률, 그리고 $Γ$는 기울기 정규화 값입니다. 이 경계는 SAM의 암묵적인 평탄성 편향을 정량적으로 설명합니다. 즉, 다른 모든 값을 고정했을 때, 더 작은 배치 크기, 더 큰 학습률 또는 더 큰 반경은 선형적으로 안정적인 SAM이 더 평탄한 최소점에 도달하도록 제한합니다. 또한, $ρ$는 평탄성을 촉진하기에는 충분히 커야 하지만, 근사 및 안정적인 훈련을 유지하기에는 충분히 로컬해야 하는 필수적인 상호 작용을 보여줍니다. 이 예측은 ResNet-18 및 VGG-19를 사용하여 CIFAR-100 데이터셋에 대해 900개의 모델을 분석한 통제된 실험에서 검증되었습니다. 여기서 $ρ$가 증가하면 배치 크기와 학습률 설정에 관계없이 최대 헤세 행렬 고유값이 감소하는 경향이 있었습니다. 마지막으로, 우리는 이 분석을 Taylor-Locality Controlled SAM (TLC-SAM)에 적용했습니다. TLC-SAM은 관찰된 Taylor 근사 오차를 사용하여 $ρ$를 조정하여, 고정 반경 SAM에 비해 최고 헤세 행렬 고유값을 더욱 줄입니다. 본 연구 결과는 SAM 변형을 분석하고 설계하기 위한 정량적인 하이퍼파라미터 경계를 제공하며, 안정성-로컬리티 관점에서 이를 설명합니다.

Original Abstract

Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. In particular, the perturbation radius $ρ$ is typically treated as an isolated tuning parameter, despite defining the neighborhood in which SAM measures sharpness. We analyze mini-batch SAM near an interpolating minimum through linear stability. Under local linearization and gradient-noise alignment assumptions, we prove that every linearly stable minimum satisfies $λ_{\max}\leq\sqrt[3]{bΓ/(2ρη^2)}$, where $λ_{\max}$ is the largest Hessian eigenvalue, $b$ is the batch size, $η$ is the learning rate, and $Γ$ bounds the gradient norm. The bound quantitatively characterizes SAM's implicit flatness bias: holding the other quantities fixed, a smaller batch size, a larger learning rate, or a larger radius restricts linearly stable SAM to flatter minima. It also exposes a necessary trade-off: $ρ$ should be large enough to promote flatness, yet remain local enough to preserve the approximation and stable training. We validate this prediction in a controlled study of 900 models on CIFAR-100 with ResNet-18 and VGG-19, where increasing $ρ$ is consistently associated with a smaller largest Hessian eigenvalue across batch-size and learning-rate settings. Finally, we instantiate the analysis in Taylor-Locality Controlled SAM (TLC-SAM), which adjusts $ρ$ using the observed Taylor-approximation error and further reduces the top Hessian eigenvalue relative to fixed-radius SAM. Our results provide quantitative hyperparameter bounds and a stability--locality perspective for analyzing and designing SAM variants.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!