2603.05995v1 Mar 06, 2026 cs.RO

TADPO: 강화 학습, 비포장 도로 환경으로 확장

TADPO: Reinforcement Learning Goes Off-road

Zhouchonghao Wu
Zhouchonghao Wu
Citations: 2
h-index: 1
Vedant Mundheda
Vedant Mundheda
Citations: 15
h-index: 2
L. Navarro-Serment
L. Navarro-Serment
Citations: 1,196
h-index: 16
C. Schoenborn
C. Schoenborn
Citations: 0
h-index: 0
Jeff Schneider
Jeff Schneider
Citations: 62
h-index: 2
R. Song
R. Song
Citations: 0
h-index: 0

비포장 도로에서의 자율 주행은 지도되지 않은, 다양한 지형과 불확실하고 복잡한 동역학을 가진 환경을 탐색해야 하므로 상당한 어려움을 야기합니다. 이러한 문제를 해결하기 위해서는 효과적인 장기 계획 수립과 적응적인 제어가 필요합니다. 강화 학습(RL)은 상호 작용을 통해 직접 제어 정책을 학습함으로써 유망한 해결책을 제시합니다. 그러나 비포장 도로 주행은 장기적인 작업이며, 보상 신호가 희소하기 때문에, 기존의 표준 RL 방법론은 이러한 환경에 적용하기 어렵습니다. 본 논문에서는 PPO(Proximal Policy Optimization)를 확장한 새로운 정책 경사(policy gradient) 방법인 TADPO를 소개합니다. TADPO는 오프라인 데이터를 활용하여 교사(teacher) 역할을 수행하고, 온라인 데이터를 활용하여 학습(student)을 진행합니다. 이러한 기반을 바탕으로, 시각 정보를 기반으로 하는 완전한 강화 학습 시스템을 개발하여 고속 비포장 도로 주행을 가능하게 하며, 극단적인 경사 및 장애물이 많은 지형을 탐색할 수 있습니다. 저희는 시뮬레이션 환경에서 성능을 검증했으며, 특히 실제 비포장 도로 차량에 대한 제로샷(zero-shot) 시뮬레이션-실제(sim-to-real) 전송을 성공적으로 수행했습니다. 저희가 알고 있는 한, 이 연구는 최초로 강화 학습 기반 정책을 실제 크기의 비포장 도로 플랫폼에 적용한 사례입니다.

Original Abstract

Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics. Addressing these challenges requires effective long-horizon planning and adaptable control. Reinforcement Learning (RL) offers a promising solution by learning control policies directly from interaction. However, because off-road driving is a long-horizon task with low-signal rewards, standard RL methods are challenging to apply in this setting. We introduce TADPO, a novel policy gradient formulation that extends Proximal Policy Optimization (PPO), leveraging off-policy trajectories for teacher guidance and on-policy trajectories for student exploration. Building on this, we develop a vision-based, end-to-end RL system for high-speed off-road driving, capable of navigating extreme slopes and obstacle-rich terrain. We demonstrate our performance in simulation and, importantly, zero-shot sim-to-real transfer on a full-scale off-road vehicle. To our knowledge, this work represents the first deployment of RL-based policies on a full-scale off-road platform.

1 Citations
0 Influential
8 Altmetric
41.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!