행동 정렬 표현을 이용한 로봇 간 이식 기술
Cross-Embodiment Transfer via Behavior-Aligned Representations
최근 로봇 조작 분야에서 대규모 모방 학습의 발전은 다양한 로봇 플랫폼에 걸친 데이터 세트를 활용하는 데 기인합니다. 그러나 상당한 로봇 간 이식 성능을 달성하는 것은 여전히 어려운 과제입니다. 본 연구에서는 시각-언어-행동(VLA) 모델에서 행동 정렬 표현(예: 객체 바운딩 박스, 언어적 동작, 로봇 운동의 엔드-이펙터 궤적)을 사용하는 것이 로봇 간 이식 성능 향상에 미치는 영향을 분석합니다. 본 연구에서는 이러한 표현 방식이 로봇 플랫폼 간 불변성을 유지하면서도 로봇 행동을 예측할 수 있다면, 대규모 데이터를 통합하여 이식을 강화하는 데 도움이 될 것이라는 가설을 제시합니다. 이를 검증하기 위해, 다양한 로봇 플랫폼에서 새로운 플랫폼으로의 이식 성능을 평가하도록 설계된 시뮬레이션 기반 벤치마크를 개발했습니다. 이 벤치마크를 사용하여 다양한 표현 방식과 통합 방법을 비교 분석했습니다. 연구 결과, 엔드-이펙터 궤적이 이식에 특히 유용하며, 일반적으로 더 큰 사전 데이터 세트가 있을 때 표현 방식의 효용성이 높아지고, 행동 정보 없이도 데이터를 활용할 수 있음을 확인했습니다. 또한, 이러한 표현 방식은 시뮬레이션에서 학습된 로봇 제어 정책을 실제 로봇으로 이전하는 과정(sim-to-real)에서의 이식 성능을 향상시키며, 실제 로봇의 작업 완료율을 28% 개선할 수 있음을 보여주었습니다. 연구 결과에 대한 자세한 내용은 다음 웹사이트에서 확인할 수 있습니다: https://ajaysridhar.com/barx/
Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: https://ajaysridhar.com/barx/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.