2602.11584v1 Feb 12, 2026 cs.LG

그래디언트 압축은 일반화를 저해할 수 있다: 합성 데이터가 유도하는 평활도 인식 최소화를 통한 해결책

Gradient Compression May Hurt Generalization: A Remedy by Synthetic Data Guided Sharpness Aware Minimization

Richeng Jin
Richeng Jin
Citations: 746
h-index: 14
Zhaoyang Zhang
Zhaoyang Zhang
Citations: 58
h-index: 3
Huaiyu Dai
Huaiyu Dai
Citations: 16
h-index: 2
Yujie Gu
Yujie Gu
Citations: 34
h-index: 3

연합 학습(FL)에서 그래디언트 압축은 성능 저하를 거의 발생시키지 않으면서 통신 효율성을 크게 향상시키는 것으로 널리 알려져 있다. 본 논문에서는 그래디언트 압축이 연합 학습, 특히 non-IID 데이터 분포 하에서 더 날카로운 손실 지형(loss landscapes)을 유발하며, 이는 일반화 능력의 저하를 시사한다는 것을 발견했다. 최근 부상하고 있는 평활도 인식 최소화(Sharpness Aware Minimization, SAM)는 널리 쓰이는 확률적 경사 하강법 적용 전에 경사 상승 단계(즉, 그래디언트로 모델에 교란을 가함)를 결합하여 평탄한 최소값(flat minima)을 효과적으로 탐색한다. 그럼에도 불구하고 연합 학습에 SAM을 직접 적용하는 것은 데이터의 이질성으로 인해 전역 교란(global perturbation)의 부정확한 추정이라는 문제를 겪는다. 기존의 접근법들은 이전 통신 라운드의 모델 업데이트를 대략적인 추정치로 활용할 것을 제안한다. 그러나 모델 업데이트에 압축이 결합될 경우 그 효과는 저해된다. 본 논문에서는 전역 모델의 궤적을 활용하여 합성 데이터를 구성하고 전역 교란의 정확한 추정을 촉진하는 FedSynSAM을 제안한다. 제안된 알고리즘의 수렴성이 입증되었으며, 그 효과를 검증하기 위해 광범위한 실험이 수행되었다.

Original Abstract

It is commonly believed that gradient compression in federated learning (FL) enjoys significant improvement in communication efficiency with negligible performance degradation. In this paper, we find that gradient compression induces sharper loss landscapes in federated learning, particularly under non-IID data distributions, which suggests hindered generalization capability. The recently emerging Sharpness Aware Minimization (SAM) effectively searches for a flat minima by incorporating a gradient ascent step (i.e., perturbing the model with gradients) before the celebrated stochastic gradient descent. Nonetheless, the direct application of SAM in FL suffers from inaccurate estimation of the global perturbation due to data heterogeneity. Existing approaches propose to utilize the model update from the previous communication round as a rough estimate. However, its effectiveness is hindered when model update compression is incorporated. In this paper, we propose FedSynSAM, which leverages the global model trajectory to construct synthetic data and facilitates an accurate estimation of the global perturbation. The convergence of the proposed algorithm is established, and extensive experiments are conducted to validate its effectiveness.

1 Citations
0 Influential
7 Altmetric
36.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!