DeepStock: 정책 정규화를 활용한 강화 학습 기반 재고 관리
DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management
심층 강화 학습(DRL)은 빅데이터와 컴퓨팅 자원을 활용하여 재고 관리 정책을 훈련하는 데 유용한 일반적인 방법론을 제공합니다. 그러나 DRL의 일반적인 구현은 훈련 과정에서 사용되는 하이퍼파라미터에 대한 높은 민감성으로 인해 종종 성공적인 결과를 얻지 못하는 경우가 많습니다. 본 논문에서는 '기본 재고'와 같은 고전적인 재고 관리 개념에 기반한 정책 정규화를 적용함으로써, 여러 DRL 방법의 하이퍼파라미터 튜닝 속도를 크게 향상시키고 최종 성능을 개선할 수 있음을 보여줍니다. 우리는 알리바바의 전자 상거래 플랫폼인 Tmall에서 정책 정규화를 적용한 DRL의 100% 배포 사례에 대한 자세한 내용을 보고합니다. 또한, 정책 정규화가 재고 관리에 가장 적합한 DRL 방법론에 대한 기존의 인식을 변화시키는 것을 보여주는 광범위한 합성 실험 결과를 포함합니다.
Deep Reinforcement Learning (DRL) provides a general-purpose methodology for training inventory policies that can leverage big data and compute. However, off-the-shelf implementations of DRL have seen mixed success, often plagued by high sensitivity to the hyperparameters used during training. In this paper, we show that by imposing policy regularizations, grounded in classical inventory concepts such as "Base Stock", we can significantly accelerate hyperparameter tuning and improve the final performance of several DRL methods. We report details from a 100% deployment of DRL with policy regularizations on Alibaba's e-commerce platform, Tmall. We also include extensive synthetic experiments, which show that policy regularizations reshape the narrative on what is the best DRL method for inventory management.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.