2602.19373v1 Feb 22, 2026 cs.LG

등방성 가우시안 표현을 통한 안정적인 심층 강화 학습

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations

Ali Saheb
Ali Saheb
Citations: 2
h-index: 1
Johan Obando-Ceron
Johan Obando-Ceron
University of Montreal/ Mila
Citations: 752
h-index: 14
Aaron C. Courville
Aaron C. Courville
Citations: 96
h-index: 3
P. Bashivan
P. Bashivan
Citations: 3,312
h-index: 16
Pablo Samuel Castro
Pablo Samuel Castro
Citations: 434
h-index: 10

심층 강화 학습 시스템은 학습 목표와 데이터 분포가 시간에 따라 변화하는 비정상성(non-stationarity)으로 인해 종종 불안정한 훈련 역학 문제를 겪는다. 우리는 비정상적 타겟 환경에서 등방성 가우시안 임베딩이 증명 가능할 정도로 유리하다는 것을 보여준다. 특히, 이는 선형 리드아웃(linear readout)에 대해 시간에 따라 변하는 타겟의 안정적인 추적을 유도하고, 고정된 분산 예산 하에서 최대 엔트로피를 달성하며, 모든 표현 차원의 균형 잡힌 사용을 촉진한다. 이 모든 요소들은 에이전트가 보다 적응적이고 안정적으로 동작할 수 있게 해준다. 이러한 통찰을 바탕으로, 우리는 훈련 중 표현(representation)을 등방성 가우시안 분포 형태로 유도하기 위해 스케치된 등방성 가우시안 정규화(Sketched Isotropic Gaussian Regularization)를 사용할 것을 제안한다. 우리는 다양한 도메인에 걸친 경험적 실험을 통해, 이 간단하고 연산 비용이 저렴한 방법이 표현 붕괴, 뉴런 휴면 및 훈련 불안정성을 감소시키는 동시에 비정상성 환경에서의 성능을 향상시킨다는 것을 입증한다.

Original Abstract

Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data distributions evolve over time. We show that under non-stationary targets, isotropic Gaussian embeddings are provably advantageous. In particular, they induce stable tracking of time-varying targets for linear readouts, achieve maximal entropy under a fixed variance budget, and encourage a balanced use of all representational dimensions--all of which enable agents to be more adaptive and stable. Building on this insight, we propose the use of Sketched Isotropic Gaussian Regularization for shaping representations toward an isotropic Gaussian distribution during training. We demonstrate empirically, over a variety of domains, that this simple and computationally inexpensive method improves performance under non-stationarity while reducing representation collapse, neuron dormancy, and training instability.

3 Citations
0 Influential
8 Altmetric
43.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!