확산 모델 기반 불확실성 인지 지연 정책 최적화
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
강화 학습은 실제 환경에서 피드백의 지연으로 인해 성능 저하를 자주 겪습니다. 기존 접근 방식은 일반적으로 관측 지연으로 인한 성능 저하를 완화하기 위해 증강된 상태를 구성하거나 실제 상태를 예측합니다. 그러나 이러한 방법은 종종 확률적 MDP에 의해 발생하는 지연된 상태와 실제 상태 간의 본질적인 불일치를 간과하는 경우가 많습니다. 우리는 그러한 불일치가 존재한다는 것을 이론적으로 증명하고, 이것이 최적 정책의 성능 저하로 이어진다는 것을 보여줍니다. 이러한 문제를 해결하기 위해, 확산 모델 기반 불확실성 인지 지연 정책 최적화 (DUPO) 방법을 제안합니다. 저희 방법은 확산 모델을 사용하여 지연된 상태 정보와 현재 상태 간의 관계를 명시적으로 모델링하고, 결과적으로 얻어지는 불일치 추정치를 활용하여 지연된 정책에 가중치를 부여합니다. 다양한 확률적 지연이 존재하는 연속 로봇 제어 작업에 대한 광범위한 실험을 통해 DUPO가 기존 방법보다 일관되게 우수한 성능을 보이며, 긴 및 임의적인 지연 시나리오에서도 효과적임을 입증했습니다.
Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing approaches typically mitigate performance degradation caused by observation delays by constructing augmented states or predicting the true states. However, these methods often overlook the inherent discrepancy between delayed state and true states induced by stochastic MDP. We theoretically prove the existence of such a discrepancy and show that it leads to the degradation of the optimal policy. To address this challenge, we propose Diffusion Guided Uncertainty Aware Delayed Policy Optimization (DUPO). Our method explicitly models the relationship between delayed state message and the current state using a diffusion model, and leverages the resulting discrepancy estimates to weight delayed policies. Extensive experiments on continuous robotic control tasks with multiple stochastic delays demonstrate that DUPO consistently outperforms existing methods and remains effective even under long and random delay scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.