X-NavDP: 그룹 Q-점수 재가중 매칭을 사용한 네비게이션 디퓨전 정책의 일반화: 새로운 행동 및 로봇 플랫폼에 대한 적용
X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching
네비게이션 디퓨전 정책의 사전 훈련은 대규모 전문가 데모 데이터를 기반으로 합니다. 이러한 데이터는 일반적으로 단일 표준 로봇에 최적화된 완전 정보를 가진 플래너를 사용하여 생성됩니다. 이는 다양한 로봇 플랫폼과 어려운 시나리오(예: 막다른 골목 탈출 또는 긴 장애물 회피)에 대한 정책의 일반화 능력을 제한합니다. 이러한 시나리오는 온보드 센서에서 얻은 제한적인 정보만으로도 다양한 즉흥적인 행동을 요구합니다. 강화 학습(RL)을 사용하여 정책을 추가 훈련하는 것은 효과적인 해결책이 될 수 있습니다. 그러나 기존 디퓨전 방식 기반의 RL 방법은 미미한 개선 효과만 가져왔습니다. 이는 디퓨전 정책의 복잡한 확률 분포로 인해 정책 경사 추정이 불안정해지고, 효율적인 정책 탐색이 어렵기 때문입니다. 이러한 문제점을 해결하기 위해, 우리는 데이터 효율성이 뛰어난 디퓨전 기반 강화 학습 후처리 프레임워크인 GQRM (Group Q-score Reweighted Matching)을 제안합니다. 우리의 프레임워크는 다음과 같은 두 가지 상호 보완적인 기능을 제공합니다: (i) 사전 훈련된 정책의 기존 지식을 유지하면서 행동을 변경하는 자체 부트스트랩 탐색 전략, 그리고 (ii) 각 상태에 대한 경로별 Q-점수를 계산하여 효율적인 재가중 점수 매칭을 수행하는 그룹 Q-점수 정규화 메커니즘. 우리는 다양한 로봇 플랫폼에서 분산 온라인 강화 학습을 수행하여 생성된 최적화된 정책인 X-NavDP는 시뮬레이션 환경에서 전체 성공률을 61.20%에서 84.28%로, 실제 환경의 어려운 경우에서는 10%에서 65%로 향상시키는 최첨단 수준의 교차 로봇 플랫폼 시각 네비게이션 성능을 달성했습니다. 코드 및 모델은 다음 링크에서 공개적으로 이용할 수 있습니다: https://yty-sky.github.io/x-navdp-project-page.
Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead ends or detouring long obstacles) that demand diverse local reactive behaviors with only onboard local observations. Post-training the policy with reinforcement learning (RL) offers a principled remedy. However, previous RL for diffusion approaches lead to only marginal improvements. This is because the intractable likelihood of diffusion policies renders policy gradients unstable in addition to inefficient policy exploration. To address these challenges, we propose a data-efficient diffusion RL post-training framework - GQRM (Group Q-score Reweighted Matching). Our framework introduces two complementary designs: (i) a self-bootstrapped exploration strategy with behavior perturbation that preserves the pretrained policy prior, and (ii) a group Q-score normalization mechanism that computes per-trajectory values on each state for efficient reweighted score matching. By conducting distributed online RL training across heterogeneous embodiments, the resulting fine-tuned policy, X-NavDP, achieves state-of-the-art cross-embodiment visual navigation performance, improving the overall success rate from 61.20% to 84.28% in simulation and 10% to 65% in real-world hard cases. The code and model are publicly available at https://yty-sky.github.io/x-navdp-project-page.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.