2606.18820v1 Jun 17, 2026 cs.LG

발전하는 마르코프 결정 과정: 증가하는 정보와 감소하는 행동 집합 하에서의 의사 결정

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

Jiaxin Liu
Jiaxin Liu
Citations: 113
h-index: 4
Zewei Dong
Zewei Dong
Citations: 2
h-index: 1
Aiping Yang
Aiping Yang
Citations: 0
h-index: 0
Shuqi Zhang
Shuqi Zhang
Citations: 0
h-index: 0
Xuebin Chen
Xuebin Chen
Citations: 0
h-index: 0
Yuhang Yang
Yuhang Yang
Citations: 376
h-index: 10
Jiangming Yang
Jiangming Yang
Citations: 0
h-index: 0

순차적 의사 결정 문제는 종종 정보의 비대칭적인 변화와 의사 결정 유연성의 변화를 보입니다. 즉, 의사 결정 주기가 진행됨에 따라 에이전트는 더 풍부한 정보를 얻는 반면, 운영상의 제한, 약속 또는 자원 제약으로 인해 실행 가능한 행동들이 사라집니다. 일반적인 MDP(마르코프 결정 과정) 모델은 이러한 구조를 단계별 상태 설명과 행동 마스크로 단순화하여, 어떤 의사 결정을 우선적으로 내려야 하는지, 어떤 것을 연기할 수 있는지 결정하는 중첩된 정보-행동 비대칭성을 가립니다. 우리는 이러한 정보-행동 비대칭성에 기반한 새로운 모델인 '발전하는 마르코프 결정 과정(MMDP)'을 제안합니다. MMDP의 주요 결과 중 하나는 '만료되는 행동 우선 순위 원칙'이며, 이는 다음 단계 전에 해결해야 할 행동들을 식별합니다. 이러한 구조에 의해 동기 부여를 받아, 우리는 단계 인지 정책 설계, 만료되는 행동 추상화, 그리고 증강 검색과 함께 학습하는 구조 인식 강화 학습 프레임워크를 개발했습니다. 제어된 다중 공급 업체 보충 문제, 복잡성이 증가하는 단순화된 현금 관리 환경 및 생산 규모 시뮬레이션에서 실험한 결과, 이러한 비대칭성을 명시적으로 모델링하면 학습 효율성을 향상시키고 의사 결정 문제가 더욱 복잡해질수록 그 가치가 높아짐을 보여줍니다.

Original Abstract

Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints. Standard MDP formulations typically flatten this structure into stage-dependent state descriptions and action masks, thereby obscuring the nested information--action asymmetry that determines which decisions are urgent and which can be deferred. We introduce Maturing Markov Decision Processes (MMDPs), a formulation built around this information--action asymmetry. We characterize one of its key consequences through an expiring-action priority principle, which identifies the actions that must be resolved before the next stage. Motivated by this structure, we develop a structure-aware reinforcement learning framework with stage-aware policy design, expiring-action abstraction, and search-augmented learning with distillation. Experiments on a controlled multi-supplier replenishment problem, simplified cash-management environments of increasing complexity, and a production-scale simulator show that explicitly modeling this asymmetry improves learning efficiency and becomes increasingly valuable as decision problems scale.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!