KAI: 데이터 효율적인 다관절 물체 조작을 위한 운동학적 정보 기반 인터페이스
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
다관절 물체 조작은 로봇 데모만으로는 학습하기 어렵고 비용이 많이 드는 운동학적 구조에 대한 이해를 필요로 합니다. 본 논문에서는 운동학적 정보를 인지하는 인터페이스(KAI)를 소개합니다. KAI는 다관절 물체의 운동학적 구조를 나타내는 체계적인 중간 표현 방식입니다. KAI는 해석 가능한 기하학적 및 운동학적 사전 지식을 정책 학습에 통합하여, 다관절 동작의 기본 구조와 일치하는 강력한 귀납적 편향을 제공합니다. 이러한 설계는 샘플 효율성을 크게 향상시키며, 특히 데이터가 부족한 환경에서 더욱 두드러진 효과를 보입니다. 여섯 가지 시뮬레이션 작업에서 KAI는 평균 성공률 82.9%를 달성하여, 데모 데이터를 절반만 사용하면서도 기존 방법과 동등하거나 더 나은 성능을 보여줍니다. 또한, 본 방법은 새로운 배경 및 시각적 방해 요소에 대한 강력한 일반화 능력을 가지며, 단일의 깨끗한 학습 환경에서 복잡한 실제 환경으로 성공적으로 전이됩니다. KAI의 액션-불문(action-agnostic) 설계는 인간과의 상호 작용 비디오를 활용한 공동 훈련을 통해 실제 환경에서의 강건성을 더욱 향상시킬 수 있습니다. 다양한 시각적 방해 요소 하에서, 비디오 공동 훈련을 적용한 본 방법은 평균 성공률 70% 이상을 달성합니다.
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-Aware Articulation Interface (KAI), a structured intermediate representation that captures the kinematic structure of articulated objects. By embedding interpretable geometric and kinematic priors into policy learning, KAI provides a strong inductive bias aligned with the underlying structure of articulated motion. This design effectively improves sample efficiency, with gains particularly pronounced in low-data regimes: across six simulation tasks, our method achieves an average success rate of 82.9%, matching or surpassing baseline performance while using only half the demonstration data. Our method also exhibits robust generalization to unseen backgrounds and visual distractors, transferring from a single clean training environment to cluttered real-world scenes. KAI's action-agnostic design further enables co-training with human interaction videos to enhance real-world robustness: under diverse visual distractions, our method with video co-training achieves over 70% average success rate.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.