2602.03778v1 Feb 03, 2026 cs.LG

L-infinity 상의 벨만 연산자를 이용한 CVaR MDP의 보상 재분배

Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity

Aneri Muni
Aneri Muni
Citations: 104
h-index: 4
Vincent Taboga
Vincent Taboga
Citations: 47
h-index: 4
Esther Derman
Esther Derman
Citations: 14
h-index: 2
Pierre-Luc Bacon
Pierre-Luc Bacon
Citations: 3,433
h-index: 22
Erick Delage
Erick Delage
Citations: 54
h-index: 3

정적 조건부 가치 위험(CVaR)과 같은 꼬리 위험 측정법은 안전이 중요한 응용 분야에서 드물지만 파괴적인 사건을 예방하기 위해 사용됩니다. 위험 중립적인 목표와 달리, 수익의 정적 CVaR은 전체 경로에 의존하며, 이는 기본 마르코프 결정 과정에서 재귀적인 벨만 분해를 허용하지 않습니다. 전통적인 해결책은 연속 변수를 사용한 상태 확장을 포함합니다. 그러나 허용 가능한 가치 함수의 특정 클래스로 제한하지 않는 한, 이러한 공식은 희소한 보상을 유발하고 퇴화된 고정점을 생성합니다. 본 연구에서는 상태 확장을 기반으로 하는 정적 CVaR 목표에 대한 새로운 공식을 제안합니다. 저희의 대안적인 접근 방식은 다음과 같은 벨만 연산자를 제공합니다: (1) 각 단계별로 밀도가 높은 보상; (2) 유한한 가치 함수 전체 공간에 대한 수축 특성. 이러한 이론적 기반을 바탕으로, 이산화된 확장된 상태를 기반으로 하는 위험 회피 가치 반복 및 모델-프리 Q-러닝 알고리즘을 개발했습니다. 또한 이산화로 인한 수렴 보장 및 근사 오차 경계를 제공합니다. 실험 결과는 저희의 알고리즘이 CVaR에 민감한 정책을 성공적으로 학습하고 효과적인 성능-안전 균형을 달성한다는 것을 보여줍니다.

Original Abstract

Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral objectives, the static CVaR of the return depends on entire trajectories without admitting a recursive Bellman decomposition in the underlying Markov decision process. A classical resolution relies on state augmentation with a continuous variable. However, unless restricted to a specialized class of admissible value functions, this formulation induces sparse rewards and degenerate fixed points. In this work, we propose a novel formulation of the static CVaR objective based on augmentation. Our alternative approach leads to a Bellman operator with: (1) dense per-step rewards; (2) contracting properties on the full space of bounded value functions. Building on this theoretical foundation, we develop risk-averse value iteration and model-free Q-learning algorithms that rely on discretized augmented states. We further provide convergence guarantees and approximation error bounds due to discretization. Empirical results demonstrate that our algorithms successfully learn CVaR-sensitive policies and achieve effective performance-safety trade-offs.

1 Citations
0 Influential
11 Altmetric
56.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!