2605.27079v1 May 26, 2026 cs.LG

신뢰 영역 Q 적절성 일치 방법 (Trust Region Q Adjoint Matching)

Trust Region Q Adjoint Matching

Changyeon Kim
Changyeon Kim
Citations: 289
h-index: 9
Kyungmin Lee
Kyungmin Lee
Citations: 317
h-index: 9
Yong Dong
Yong Dong
Citations: 7
h-index: 2
Jinwoo Shin
Jinwoo Shin
Citations: 876
h-index: 10
Jaehyuk Kim
Jaehyuk Kim
Citations: 78
h-index: 2

사전 학습된 정책을 사용하는 오프라인 강화 학습은 다단계 샘플링 과정에서 발생하는 최적화 불안정성으로 인해 여전히 어려운 과제입니다. 최근에는 Q-러닝과 적절성 일치 방법(QAM)이 학습된 평가기(critic)를 사용하여 메모리가 없는 확률적 최적 제어(SOC) 문제로 재구성함으로써 이러한 문제를 해결하고자 했습니다. 그러나 QAM은 평가기 기반 개선의 근본적인 취약성을 상속합니다. 즉, 평가기의 작은 오류가 평가기가 불안정할 경우 증폭되어 모델 붕괴를 초래하는 경우가 많습니다. 본 논문에서는 신뢰 영역 Q-적절성 일치 방법(TRQAM)을 소개합니다. TRQAM은 사전 학습된 정책을 사용한 경로 공간 KL 발산(path-space KL divergence)을 적응적으로 제어하는 안정적인 오프라인 미세 조정 알고리즘이며, 투영 이중 하강법(projected dual descent)을 통해 구현됩니다. 특히, SOC 동역학에서 신뢰 영역 매개변수 λ를 최적화하며, 경로 공간 KL 발산이 λ의 닫힌 형태 함수로 표현될 수 있음을 이론적으로 증명합니다. 결과적으로, TRQAM은 사전 학습된 정책으로부터 정확하게 벗어나는 정도를 정밀하게 제어하여 안정적인 오프라인 강화 학습을 달성할 수 있습니다. 50개의 OGBench 작업에 대한 실험에서, TRQAM은 오프라인 강화 학습 및 오프라인-온라인 강화 학습 모두에서 기존 방법보다 일관되게 우수한 성능을 보였습니다. 특히, TRQAM은 오프라인 강화 학습에서 68%의 성공률을 달성하여, 가장 강력한 기준 모델인 46%를 크게 능가했습니다.

Original Abstract

Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this issue by reformulating into a memoryless stochastic optimal control (SOC) problem with a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement: small critic errors are amplified when critics are ill-conditioned, often leading to model collapse. This paper introduces Trust Region Q-Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL with pretrained flow policies through projected dual descent. Specifically, we optimize the trust-region parameter $λ$ in SOC dynamics, and theoretically show that the path-space KL can be represented by a closed-form function of $λ$. As a result, our method can precisely control the exact deviation from pretrained flow policies, achieving stable off-policy RL. Through experiments on 50 OGBench tasks, TRQAM consistently outperforms prior arts in both offline RL and offline-to-online RL. In particular, TRQAM achieves an overall success rate of 68% in offline RL, substantially improves the strongest baseline at 46%.

1 Citations
0 Influential
5 Altmetric
26.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!