모멘트 클로저를 이용한 불확실성 하에서의 분석적 계획
Analytic Planning under Uncertainty with Moment Closure
확률적인 환경에서 효과적인 모델 기반 강화 학습은 예측 불확실성을 고려하는 계획을 필요로 합니다. 전체 상태 분포를 분석적으로 전파하는 것은 이러한 작업을 수행하는 원칙적인 방법이지만, 일반적으로 계산 가능하도록 제한적인 정책 또는 보상 구조가 요구되어 왔습니다. 결과적으로, 현대 딥 강화 학습은 상당 부분 확률적 샘플링(이는 상당한 목표 분산을 유발함) 또는 예측 공분산을 완전히 무시하는 결정론적 점 추정으로 회귀했습니다. 본 연구에서는 이러한 제약 조건 없이 분포 인지 계획이 가능한지 조사합니다. 이차적인 행동-값 매개변수화를 사용하여, 먼저 벨만 업데이트를 상태-값 함수에 대한 기댓값으로 줄입니다. 핵심 아이디어는 예측 전환 분포와 값 함수 클래스 간의 호환성 원칙이며, 이 원칙 하에서 해당 기댓값은 분포의 모멘트에 대해 분석적으로 표현됩니다. 본 연구에서는 가우시안 전환 모델과 방사형 기반 값 함수를 결합하여 이 원칙을 구현하고, 예측 평균과 공분수를 모두 전파하는 닫힌 형태의 업데이트를 얻습니다. 실험 결과, 제안된 방법은 연속적인 제어 환경에서 확률적 관찰 하에 목표 분산을 줄이고 잘 조정된 예측 불확실성을 제공하며, 학습된 분포 모델을 사용한 계획을 위한 원칙적인 프레임워크를 제시합니다.
Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.