오프라인 강화 학습에서 국소 수정 전파 제어를 위한 볼록 폐포 이웃 스무딩 이중 일반화
Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL
오프라인 강화 학습(offline RL)은 주변의 분포 외 액션(out-of-distribution, OOD)으로부터 이점을 얻을 수 있지만, 이러한 액션에서의 추정 오류는 부트스트래핑에 의해 증폭될 수 있습니다. 기존의 정규화 및 로컬 일반화 방법은 허용 가능한 OOD 영역 또는 일반화된 목표의 영향 중 하나를 제어하는 경향이 있으며, 종종 별도의 메커니즘을 통해 이루어집니다. 본 논문에서는 벨만 방정식의 업데이트를 샘플 내 값 목표와 CHN(Convex Hull Neighborhood)-로컬 수정 항의 합으로 표현하는 Convex Hull Neighborhood Smooth Dual Generalization (CSDG) 기법을 제안합니다. 이러한 수식은 일반화된 기여도를 명시적으로 나타내고, 샘플 내 참조 경로와 분리합니다. 수정항은 서로 다른 섭동 반경에서 샘플링된 샘플 내 지향 및 OOD 지향 후보들을 스무딩하여 얻어집니다. 혼합 계수 lambda는 각 업데이트에 대한 이 항의 기여도를 조절하며, 할인율은 gamma로 유지됩니다. 경계 조건과 고정된 섭동 커널 하에서, 우리는 정확한 단일 단계 수정 식, 시간 변화형 반복 오차 한계, 그리고 고정점에서의 분기 불일치에만 의존하는 고정점 오차 한계를 유도했습니다. 또한 이상적인 연산자에 의해 유도되는 암시적 정책을 특성화하고 조건부 비열화 기준을 제시합니다. 실제 알고리즘은 비대칭 경계 노이즈와 기대값 회귀를 사용하여 이러한 양들을 근사하며, 정확한 지지 집합 분류나 추가적인 비관적인 OOD 페널티가 필요하지 않습니다. Gym-MuJoCo 및 AntMaze 환경에서의 실험 결과는 우수한 성능과 안정적인 값 추정 결과를 보여줍니다. 코드: https://github.com/YOUNG-fnxm/CSDG
Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. Existing regularization and local-generalization methods control either the admissible OOD region or the influence of generalized targets, often through separate mechanisms. We propose Convex Hull Neighborhood Smooth Dual Generalization (CSDG), which expresses the Bellman backup as an in-sample value target plus a CHN-local correction. This formulation makes the generalized contribution explicit and separates it from the in-sample reference path. The correction is obtained by smoothing in-sample-oriented and OOD-oriented candidates sampled at different perturbation radii. A mixture coefficient lambda scales its contribution to each backup, while the recursive discount remains gamma. Under boundedness and fixed perturbation kernels, we derive an exact one-step correction identity, a time-varying iterate bound, and a fixed-point bound that depends only on the branch discrepancy at the fixed point. We further characterize the implicit policies induced by the idealized operators and give a conditional non-degradation criterion. The practical algorithm approximates these quantities using asymmetric bounded noise and expectile regression, without exact support classification or an additional pessimistic OOD penalty. Experiments on Gym-MuJoCo and AntMaze show strong aggregate performance and stable value estimation. Code is available at: https://github.com/YOUNG-fnxm/CSDG
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.