2606.16480v1 Jun 15, 2026 cs.RO

HOLOMPPI: 계층적 정책 최적화를 통한 다중 시나리오 동작 계획

HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization

Faizan M. Tariq
Faizan M. Tariq
Citations: 102
h-index: 6
Sangjae Bae
Sangjae Bae
Citations: 138
h-index: 8
David Isele
David Isele
Citations: 1,911
h-index: 17
Youngjae Min
Youngjae Min
Citations: 51
h-index: 2
Jovin D'sa
Jovin D'sa
Citations: 97
h-index: 6
Navid Azizan
Navid Azizan
Massachusetts Institute of Technology
Citations: 1,561
h-index: 18

실제 환경에 배치되는 로봇은 각 시나리오별로 별도의 조정 없이 다양한 시나리오에서 동작을 계획해야 합니다. 엔드-투-엔드 강화 학습(RL)은 여러 시나리오에서 일반화 가능하지만, 데이터 분포 변화, 보상 오류 및 확률적 상호 작용으로 인해 종종 불안정해지는 경향이 있습니다. 모델 예측 경로 적분(MPPI) 제어는 그래디언트 없이 강력한 실시간 개선을 제공하지만, 성능은 잘 정의된 샘플링 사전 지식에 의존하며, 수동으로 사전 지식을 설계하는 것은 다중 시나리오 배포에는 적합하지 않습니다. 본 논문에서는 고수준 정책 학습과 저수준 확률적 최적 제어를 결합한 다중 시나리오 동작 계획 프레임워크인 HOLOMPPI(High-level Offline, Low-level Online MPPI)를 제시합니다. 오프라인 단계에서는 추상적인 행동 공간에서 시나리오에 강건한 계획을 제안하는 고수준 정책을 학습하고, 온라인 로우롤아웃을 위한 학습된 세계 모델을 사용합니다. 온라인 단계에서는 이 정책이 데이터 기반의 사전 지식 생성기로 작동하며, 현재 관찰 및 목표에 따라 MPPI의 샘플링 분포를 매개변수화합니다. MPPI는 그런 다음 실시간으로 이 사전 지식을 중심으로 저수준 제어 시퀀스를 최적화하여 국소적인 방해 요인에 적응합니다. 본 논문에서는 효과적인 고수준 행동 공간과 맞춤형 모델 아키텍처를 설계하여 자율 주행 시스템에 HOLOMPPI를 구현했습니다. 다양한 주행 시나리오에서의 평가 결과, HOLOMPPI는 MPPI 및 엔드-투-엔드 RL의 기존 방법보다 성능이 향상되었으며, 실시간 제어 능력을 유지하는 것을 확인했습니다.

Original Abstract

Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning. End-to-end reinforcement learning (RL) can generalize across scenarios but often becomes brittle under distribution shift, reward misspecification, and stochastic interactions. Model predictive path integral (MPPI) control enables strong real-time refinement without gradients, but its performance depends on a well-shaped sampling prior, while manually designing the priors does not scale to multi-scenario deployment. We present HOLO-MPPI (High-level Offline, Low-level Online MPPI), a multi-scenario motion planning framework that combines high-level policy learning with low-level stochastic optimal control. Offline, we learn a high-level policy that proposes scenario-robust plans in an abstract action space, with a learned world model for online rollout. Online, the policy serves as a data-driven prior generator that parameterizes MPPI's sampling distribution conditioned on the current observation and goal. MPPI then optimizes low-level control sequences around this prior in real time to adapt to local disturbances. We instantiate HOLO-MPPI in autonomous driving by designing an effective high-level action space and tailored model architectures. Our evaluation across diverse driving scenarios shows that HOLO-MPPI improves upon MPPI and end-to-end RL baselines while maintaining real-time control.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!