2608.03108v1 Aug 04, 2026 cs.LG

오프라인 강화 학습에서 국소 수정 전파 제어를 위한 볼록 폐포 이웃 스무딩 이중 일반화

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL

Zhennan Chen
Zhennan Chen
Citations: 96
h-index: 5
Yi Yang
Yi Yang
Citations: 143
h-index: 6
Mingfeng Lv
Mingfeng Lv
Citations: 0
h-index: 0
Hanlei Li
Hanlei Li
Citations: 1
h-index: 1
Zhengsen Ruan
Zhengsen Ruan
Citations: 6
h-index: 2
Lvqing Yang
Lvqing Yang
Citations: 29
h-index: 3

오프라인 강화 학습(offline RL)은 주변의 분포 외 액션(out-of-distribution, OOD)으로부터 이점을 얻을 수 있지만, 이러한 액션에서의 추정 오류는 부트스트래핑에 의해 증폭될 수 있습니다. 기존의 정규화 및 로컬 일반화 방법은 허용 가능한 OOD 영역 또는 일반화된 목표의 영향 중 하나를 제어하는 경향이 있으며, 종종 별도의 메커니즘을 통해 이루어집니다. 본 논문에서는 벨만 방정식의 업데이트를 샘플 내 값 목표와 CHN(Convex Hull Neighborhood)-로컬 수정 항의 합으로 표현하는 Convex Hull Neighborhood Smooth Dual Generalization (CSDG) 기법을 제안합니다. 이러한 수식은 일반화된 기여도를 명시적으로 나타내고, 샘플 내 참조 경로와 분리합니다. 수정항은 서로 다른 섭동 반경에서 샘플링된 샘플 내 지향 및 OOD 지향 후보들을 스무딩하여 얻어집니다. 혼합 계수 lambda는 각 업데이트에 대한 이 항의 기여도를 조절하며, 할인율은 gamma로 유지됩니다. 경계 조건과 고정된 섭동 커널 하에서, 우리는 정확한 단일 단계 수정 식, 시간 변화형 반복 오차 한계, 그리고 고정점에서의 분기 불일치에만 의존하는 고정점 오차 한계를 유도했습니다. 또한 이상적인 연산자에 의해 유도되는 암시적 정책을 특성화하고 조건부 비열화 기준을 제시합니다. 실제 알고리즘은 비대칭 경계 노이즈와 기대값 회귀를 사용하여 이러한 양들을 근사하며, 정확한 지지 집합 분류나 추가적인 비관적인 OOD 페널티가 필요하지 않습니다. Gym-MuJoCo 및 AntMaze 환경에서의 실험 결과는 우수한 성능과 안정적인 값 추정 결과를 보여줍니다. 코드: https://github.com/YOUNG-fnxm/CSDG

Original Abstract

Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be amplified by bootstrapping. Existing regularization and local-generalization methods control either the admissible OOD region or the influence of generalized targets, often through separate mechanisms. We propose Convex Hull Neighborhood Smooth Dual Generalization (CSDG), which expresses the Bellman backup as an in-sample value target plus a CHN-local correction. This formulation makes the generalized contribution explicit and separates it from the in-sample reference path. The correction is obtained by smoothing in-sample-oriented and OOD-oriented candidates sampled at different perturbation radii. A mixture coefficient lambda scales its contribution to each backup, while the recursive discount remains gamma. Under boundedness and fixed perturbation kernels, we derive an exact one-step correction identity, a time-varying iterate bound, and a fixed-point bound that depends only on the branch discrepancy at the fixed point. We further characterize the implicit policies induced by the idealized operators and give a conditional non-degradation criterion. The practical algorithm approximates these quantities using asymmetric bounded noise and expectile regression, without exact support classification or an additional pessimistic OOD penalty. Experiments on Gym-MuJoCo and AntMaze show strong aggregate performance and stable value estimation. Code is available at: https://github.com/YOUNG-fnxm/CSDG

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!