2603.19621v1 Mar 20, 2026 cs.LG

DeepStock: 정책 정규화를 활용한 강화 학습 기반 재고 관리

DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management

Yaqi Xie
Yaqi Xie
Citations: 13
h-index: 2
Xinru Hao
Xinru Hao
Citations: 2
h-index: 1
Jiaxin Liu
Jiaxin Liu
Citations: 113
h-index: 4
Will Ma
Will Ma
Citations: 19
h-index: 2
Linwei Xin
Linwei Xin
Citations: 19
h-index: 2
Lei Cao
Lei Cao
Citations: 263
h-index: 10
Yidong Zhang
Yidong Zhang
Citations: 678
h-index: 10

심층 강화 학습(DRL)은 빅데이터와 컴퓨팅 자원을 활용하여 재고 관리 정책을 훈련하는 데 유용한 일반적인 방법론을 제공합니다. 그러나 DRL의 일반적인 구현은 훈련 과정에서 사용되는 하이퍼파라미터에 대한 높은 민감성으로 인해 종종 성공적인 결과를 얻지 못하는 경우가 많습니다. 본 논문에서는 '기본 재고'와 같은 고전적인 재고 관리 개념에 기반한 정책 정규화를 적용함으로써, 여러 DRL 방법의 하이퍼파라미터 튜닝 속도를 크게 향상시키고 최종 성능을 개선할 수 있음을 보여줍니다. 우리는 알리바바의 전자 상거래 플랫폼인 Tmall에서 정책 정규화를 적용한 DRL의 100% 배포 사례에 대한 자세한 내용을 보고합니다. 또한, 정책 정규화가 재고 관리에 가장 적합한 DRL 방법론에 대한 기존의 인식을 변화시키는 것을 보여주는 광범위한 합성 실험 결과를 포함합니다.

Original Abstract

Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.

3 Citations
0 Influential
5 Altmetric
28.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!