2605.26478v1 May 26, 2026 cs.RO

확률적 분리 정책 경사법을 이용한 효율적인 온라인 시각 강화 학습

Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient

Daniel Rakita
Daniel Rakita
Citations: 12
h-index: 3
Haoxiang You
Haoxiang You
Citations: 8
h-index: 1
Davis Zong
Davis Zong
Citations: 0
h-index: 0
Teeratham Vitchutripop
Teeratham Vitchutripop
Citations: 0
h-index: 0
Ian Abraham
Ian Abraham
Citations: 6
h-index: 1
Yilang Liu
Yilang Liu
Citations: 1,169
h-index: 17
Qian Wang
Qian Wang
Citations: 1
h-index: 1
Qi Wang
Qi Wang
Citations: 31
h-index: 2

본 논문에서는 가벼운 시각 강화 학습(RL) 방법인 확률적 분리 정책 경사법(SDPG)을 제시합니다. SDPG는 단일 NVIDIA RTX 4080 GPU에서 몇 시간 이내에 다양한 시각-운동 제어 정책을 종단 간 결합 방식으로 학습할 수 있습니다. SDPG는 경로 추론의 무작위 변동을 통해 정책 경사를 추정하며, 기존 방법에 비해 훨씬 적은 양의 배치 렌더링 환경을 필요로 하여 계산 및 메모리 오버헤드를 크게 줄입니다. 시각 MuJoCo 벤치마크에서 SDPG는 학습 시간, 메모리 사용량 및 보상 측면에서 기존 방법보다 일관되게 우수한 성능을 보였습니다. 또한, 향후 연구를 지원하기 위해 정교한 조작, 어려운 이동 등 다양한 현실적인 시각 로봇 벤치마크 세트를 소개하고, 실제 하드웨어에서의 효과적인 시뮬레이션-실제 전송 결과를 보여줍니다.

Original Abstract

We present the stochastic decoupled policy gradient (SDPG), a lightweight visual reinforcement learning (RL) method that trains diverse visuomotor control policies end-to-end within a few hours on a single NVIDIA RTX 4080 GPU. SDPG estimates policy gradients via random perturbations of trajectory rollouts, requiring orders of magnitude fewer batch-rendered environments and substantially reducing compute and memory overhead. On visual MuJoCo benchmarks, SDPG consistently outperforms baseline methods in training time, memory usage, and rewards. Finally, to support future research, we introduce a suite of realistic visual robotics benchmarks spanning dexterous manipulation, challenging locomotion, and demonstrate effective sim-to-real transfer on physical hardware.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!