2606.25797v1 Jun 24, 2026 cs.AI

마르코프 결정 프로세스 온라인 통계 모델 검증을 위한 신뢰 구간

Confidence Sequences for Online Statistical Model Checking of Markov Decision Processes

Tobias Meggendorfer
Tobias Meggendorfer
Citations: 470
h-index: 11
Konstantin Kueffner
Konstantin Kueffner
Citations: 64
h-index: 4
Maximilian Weininger
Maximilian Weininger
Citations: 659
h-index: 12
Patrick Wienhoft
Patrick Wienhoft
Citations: 0
h-index: 0

마르코프 결정 프로세스(MDP)는 불확실성 하에서의 의사 결정을 모델링하는 고전적인 방법으로, 비결정적 선택과 확률적 불확실성을 모두 포함합니다. 전통적으로는 기본 확률에 대한 정확한 지식이 있다고 가정하지만, 이는 사이버 물리 시스템이나 생물학적 프로세스를 모델링할 때 종종 비현실적입니다. 여기서 통계적 방법은 의미 있는 보장을 얻을 수 있는 방법을 제공합니다. 일반적인 접근 방식은 MDP에서 샘플을 수집하고, 이러한 샘플을 사용하여 전이 확률에 대한 통계적 결론을 도출하고, 이를 통해 실제 값에 대한 경계를 얻는 것입니다. 그런 다음 이러한 경계가 너무 넓으면 다시 반복합니다. 그러나 기존의 이 접근 방식 구현은 미묘하게 부정확하거나 최적이 아니며, 종종 둘 다입니다. 우리는 특정 '온라인' 환경을 위해 설계된 여러 가지 '신뢰 구간'을 제시하고, 이를 모두 효율적인 도구에 구현하며, 그 실용성을 보여줍니다. 특히, 제안하는 방법이 기존의 '합집합 경계' 방식보다 성능이 뛰어나며, 전반적으로 우리의 구현은 평균적으로 이전 최고 수준의 방법에 비해 50배 적은 샘플을 필요로 합니다.

Original Abstract

Markov decision processes (MDPs) are a classic model of decision making under uncertainty, exhibiting both non-deterministic choice as well as probabilistic uncertainty. Traditionally, exact knowledge of the underlying probabilities is assumed. However, this often is unrealistic, e.g.\ when modelling cyber-physical systems or biological processes. Here, statistical methods provide a way towards obtaining meaningful guarantees. The classical approach is to gather samples in the MDP, use these to draw statistical conclusions about the transition probabilities, and from there obtain bounds on the true value; then, if these bounds are too broad, repeat. However, existing implementations of this approach are either subtly incorrect or sub-optimal, and quite often both. We present several \emph{confidence sequences}, which are specifically designed for such \enquote{online} settings, implement all of them in an efficient tool, and show their practical applicability. In particular, we show that they outperform classical \enquote{union-bound} style approaches, and overall our implementation requires 50x less samples on average than previous state of the art.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!