속성 기반 인과 추상화를 이용한 마르코프 결정 과정
Property-driven Causal Abstractions for Markov Decision Processes
마르코프 결정 과정(MDP)은 의사결정 모델로 널리 사용되며, 주로 상태 변수와 그 값들을 통해 구조화된 상태 공간에서 정의됩니다. 상태의 수가 기하급수적으로 증가하면 MDP에서의 많은 추론 작업이 어려워집니다. 추상화는 MDP를 줄이고 확장성 문제를 완화하는 유망한 기술입니다. 본 연구에서는 구조화된 MDP에 대한 인과 관계 개념을 도입하고, 원래 MDP 모델의 많은 특성을 유지하는 새로운 속성 기반 인과 추상화 기법을 제시합니다. 이를 위해 상태 변수 술어에 대한 인과 관계를 활용하고, 주어진 추상화 속성을 만족하거나 위반하는 이유가 동일한 상태들을 식별합니다. 우리는 MDP, 구간 MDP, 확률 게임 등 다양한 모델 유형을 사용하여 다양한 인과 MDP 추상화를 이론적 및 실증적으로 비교합니다. 우리의 평가는 다음과 같은 잠재력을 보여줍니다: 여러 표준 벤치마크에서 작은 추상화를 얻어 원래 MDP에 대한 거의 최적의 정책을 계산할 수 있습니다. 또한, 우리의 인과 추상화는 종종 관련 있는 대규모 MDP 모델로 일반화됩니다.
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.