2606.29820v1 Jun 29, 2026 cs.LG

상태 인지 탐색을 통한 이중 흐름 강화 학습

Dual-Flow Reinforcement Learning with State-Aware Exploration

Qijun Li
Qijun Li
Citations: 0
h-index: 0
Zheng Fu
Zheng Fu
Citations: 291
h-index: 10
Qi Song
Qi Song
Citations: 18
h-index: 2
Yifei He
Yifei He
Citations: 106
h-index: 4
Weitao Zhou
Weitao Zhou
Citations: 250
h-index: 7
Kun Jiang
Kun Jiang
Citations: 1
h-index: 1
Diange Yang
Diange Yang
Citations: 491
h-index: 12

복잡한 연속 제어 강화 학습 작업에서 최적의 행동은 종종 불확실하고 다중 모드를 갖는 보상 분포와 함께 나타나는데, 이는 신뢰할 수 있는 가치 추정 및 다중 모드 탐색을 어렵게 만듭니다. 단일 모드 가우시안을 사용하는 기존 가치 추정 방법은 표현력을 제한하며 편향된 결과를 초래합니다. 최근 생성 정책은 다중 모드 행동을 표현할 수 있지만, 종종 몇 가지 모드로 붕괴되고 높은 가치를 갖는 행동 공간 영역을 충분히 탐색하지 못하는 경향이 있습니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 조건부 흐름 매칭(CFM)을 사용하여 연속적인 보상 분포와 다중 모드 정책 분포를 동시에 모델링하는 통합된 액터-크리틱 프레임워크인 Dual-Flow RL을 제안합니다. 이 설계는 신뢰할 수 있는 가치 추정과 지속적인 다중 모드 탐색을 지원합니다. 또한, 정책 엔트로피 및 행동 불확실성 공분산을 활용하여 상태에 따른 탐색 조절을 가능하게 하는 Entropy-Covariance Exploration Regulator (ECER)를 도입하여 탐색 성능을 더욱 향상시킵니다. DeepMind Control Suite와 Humanoid-Bench에서의 실험 결과, Dual-Flow RL은 대부분의 작업에서 최첨단 성능을 달성했으며, 기존 확산 기반 및 흐름 기반 방법보다 훨씬 우수한 성능을 보였습니다.

Original Abstract

In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging. Existing value estimation methods using unimodal Gaussians restrict expressiveness and yield biased estimates. Recent generative policies can represent multimodal actions but often collapse to a few modes and under-explore high-value areas of the action space. Motivated by these challenges, we propose Dual-Flow RL, a unified actor-critic framework that jointly models a continuous return distribution and a multimodal policy distribution using conditional flow matching (CFM). This design supports reliable value estimation and sustained multimodal exploration. To further enhance exploration, we introduce an Entropy-Covariance Exploration Regulator (ECER) that enables state-aware exploration regulation leveraging policy entropy and action-uncertainty covariance. Experiments on DeepMind Control Suite and Humanoid-Bench show that Dual-Flow RL achieves state-of-the-art performance on most tasks, significantly outperforming prior diffusion-based and flow-based methods.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!