주변 확산 정책: 로봇 공학 분야의 하위 최적 데이터로부터의 모방 학습
Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics
본 논문에서는 로봇 공학 분야에서 하위 최적 데이터를 활용한 모방 학습을 위한 간단하고 체계적인 방법인 '주변 확산 정책(Ambient Diffusion Policy)'을 제안합니다. 고품질의 작업별 로봇 데이터는 수집하는 데 비용과 시간이 많이 소요되는 반면, 품질이 낮거나 분포가 다른 하위 최적 데이터 세트는 풍부하게 존재합니다. 기존의 로봇 공학 분야에서 두 가지 데이터 소스를 함께 학습시키는 방법은 종종 하위 최적 샘플에 존재하는 유용한 특징과 해로운 특징을 분리하는 데 실패합니다. 반면, 본 연구에서는 새로운 접근 방식을 도입하여 로봇 공학 분야의 공동 훈련 과정에서 '노이즈 의존적인 데이터 활용'이라는 새로운 축을 제시함으로써, 오직 유용한 특징만을 추출합니다. 주변 확산 정책은 훈련 과정에서 하위 최적 데이터가 기여하는 범위를 높은 확산 시간과 낮은 확산 시간으로 제한합니다. 본 연구의 접근 방식을 엄밀하게 뒷받침하기 위해, 먼저 로봇 액션 데이터가 스펙트럼 거듭제곱 법칙을 따른다는 것을 관찰하고, 이를 바탕으로 최적의 확산 정책이 갖는 두 가지 중요한 특징인 '전역-로컬 계층 구조'와 '지역성'을 활용합니다. 이러한 내용을 간략화된 모델을 사용하여 이론적으로 공식화했습니다. 실험 결과, 제안하는 주변 확산 정책은 노이즈가 많은 경로, 시뮬레이션-실제 격차, 작업 불일치, 대규모 데이터 혼합 등 다양한 유형의 하위 최적 액션 데이터에 대해 긍정적인 성능을 보였습니다. 특히, Open X-Embodiment와 같이 이질적인 품질과 비정형적인 분포 변화를 가진 대규모 데이터 세트에 적용했을 때, 기존의 공동 학습 방법보다 최대 33% 향상된 성능을 보여주었습니다. 전반적으로, 주변 확산 정책은 하위 최적 데모의 활용도를 높이고 로봇 공학 분야에서 사용 가능한 데이터 소스의 범위를 확장합니다.
We propose Ambient Diffusion Policy, a simple and principled method for imitation learning from suboptimal data in robotics. High-quality, task-specific robot data is expensive and time-consuming to collect, while suboptimal datasets with lower-quality or out-of-distribution demonstrations are abundant. Existing methods that co-train on both data sources in robotics often fail to separate the meaningful and the harmful features in the suboptimal samples. In contrast, our method extracts only the useful features by introducing a new axis to co-training in robotics: noise-dependent data usage. Ambient Diffusion Policy restricts the contribution of suboptimal data during training to only the high and low diffusion times. To rigorously justify our approach, we first observe that robot action data exhibits a spectral power law. This induces two important properties on the optimal Diffusion Policy that we exploit: a global-to-local hierarchy and locality. We theoretically formalize this discussion using a simplified model. Our experiments validate Ambient Diffusion Policy on four types of suboptimal action data (noisy trajectories, sim-to-real gap, task mismatch, and large-scale data mixtures) across six tasks. The results show that it effectively learns from arbitrary sources of suboptimal data. Notably, it outperforms existing co-training baselines by up to 33% when scaled to Open X-Embodiment - a large dataset with heterogeneous data quality and unstructured distribution shifts. Overall, Ambient Diffusion Policy increases the utility of suboptimal demonstrations and expands the set of usable data sources in robotics.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.