2601.18420v1 Jan 26, 2026 cs.LG

경사 기반 정규화 자연 경사

Gradient Regularized Natural Gradients

Satya Dash
Satya Dash
Citations: 2,292
h-index: 20
Samuel Kaski
Samuel Kaski
Citations: 55
h-index: 2
Mingfei Sun
Mingfei Sun
Citations: 44
h-index: 2
Hossein Abdi
Hossein Abdi
Citations: 3
h-index: 1
Wei Pan
Wei Pan
Citations: 20
h-index: 2

경사 기반 정규화(GR)는 학습된 모델의 일반화 성능을 향상시키는 것으로 나타났습니다. 자연 경사 하강법은 학습 초기 단계에서 최적화 속도를 가속화하는 것으로 알려져 있지만, 2차 최적화기의 학습 역학이 GR로부터 어떻게 이점을 얻을 수 있는지에 대한 연구는 상대적으로 부족했습니다. 본 연구에서는 명시적인 경사 정규화를 자연 경사 업데이트와 통합하는 확장 가능한 2차 최적화기인 Gradient-Regularized Natural Gradients (GRNG)를 제안합니다. 우리의 프레임워크는 두 가지 상호 보완적인 알고리즘을 제공합니다. 첫 번째는 구조화된 근사를 통해 Fisher 정보 행렬(FIM)의 명시적인 역산을 피하는 빈도주의적 변형이며, 두 번째는 정규화된 칼만(Regularized-Kalman) 수식을 기반으로 하며, FIM 역산의 필요성을 완전히 제거합니다. 우리는 GRNG에 대한 수렴 보장을 확립하여, 경사 정규화가 안정성을 향상시키고 전역 최소값으로의 수렴을 가능하게 함을 보여줍니다. 실험적으로, GRNG는 1차 방법(SGD, AdamW) 및 2차 기준(K-FAC, Sophia)과 비교하여 일관되게 최적화 속도와 일반화 성능을 향상시켰으며, 컴퓨터 비전 및 자연어 처리 벤치마크에서 뛰어난 결과를 보였습니다. 우리의 연구 결과는 경사 정규화를 대규모 딥러닝을 위한 자연 경사 방법의 견고성을 향상시키는 원칙적이고 실용적인 도구로 강조합니다.

Original Abstract

Gradient regularization (GR) has been shown to improve the generalizability of trained models. While Natural Gradient Descent has been shown to accelerate optimization in the initial phase of training, little attention has been paid to how the training dynamics of second-order optimizers can benefit from GR. In this work, we propose Gradient-Regularized Natural Gradients (GRNG), a family of scalable second-order optimizers that integrate explicit gradient regularization with natural gradient updates. Our framework provides two complementary algorithms: a frequentist variant that avoids explicit inversion of the Fisher Information Matrix (FIM) via structured approximations, and a Bayesian variant based on a Regularized-Kalman formulation that eliminates the need for FIM inversion entirely. We establish convergence guarantees for GRNG, showing that gradient regularization improves stability and enables convergence to global minima. Empirically, we demonstrate that GRNG consistently enhances both optimization speed and generalization compared to first-order methods (SGD, AdamW) and second-order baselines (K-FAC, Sophia), with strong results on vision and language benchmarks. Our findings highlight gradient regularization as a principled and practical tool to unlock the robustness of natural gradient methods for large-scale deep learning.

0 Citations
0 Influential
10 Altmetric
50.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!