LUCID: 비체(embodiment)에 독립적인 의도 모델 학습 - 인터넷 기반의 비정형 인간 영상 데이터를 활용하여 로봇 기술 습득을 위한 확장 가능한 방법
LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition
현재 가장 널리 사용되는 로봇 학습 시스템은 주로 로봇 데모 또는 정형화된 인간 데이터로부터 기술을 학습하는데, 이는 수집 비용이 많이 들고 특정 로봇 플랫폼에 종속적이라는 단점이 있습니다. 반면, 비정형 인간 영상은 확장 가능한 대안을 제공합니다. 이러한 영상에는 다양한 물체, 환경 및 전략에 대한 조작 시연이 포함되어 있지만, 직접적으로 로봇 동작과 연결되지는 않습니다. 본 논문에서는 인터넷 규모의 데이터셋에서 얻은 비정형 인간 영상으로부터 작업 의도를 학습하고, 대규모 병렬 시뮬레이션 환경에서 로봇 제어를 학습하는 두 단계 프레임워크인 LUCID를 제안합니다. 의도 모델은 현재 관찰 정보를 기반으로 단기적인 의도를 예측하며, 이는 폐루프 시스템에서 작동합니다. 특정 로봇 플랫폼에 맞는 센서-운동 정책은 이 의도를 로봇 동작으로 변환합니다. 이러한 의도 인터페이스는 여러 컨트롤러에서 공유되므로, 동일한 의도 모델을 사용하여 주된 숙련형 핸드부터 병렬 턱 그리퍼와 같은 다양한 로봇 플랫폼에 적용할 수 있습니다. 우리는 LUCID를 다섯 가지 실제 조작 작업(교반, 닦기, 분류)에 대해 평가했으며, 이 작업들은 인터넷 영상만을 사용하여 감독되었으며, 새로운 환경과 물체 인스턴스로의 제로샷 전송이 가능했습니다. 또한, push-T 및 케이블 라우팅 작업은 각각 1시간씩 자체 수집한 스마트폰 영상을 사용하여 학습되었습니다. 프로젝트 페이지: https://lucid-robot.github.io/.
The most widely-adopted robot learning pipelines today learn skills from robot demonstrations or structured human data, which are expensive to collect and tied to specific embodiments. In contrast, unstructured human videos provide a scalable alternative. They contain diverse manipulation demonstrations across objects, scenes, and strategies, but are not directly connected to robot action. We propose LUCID, a two-stage framework that learns task intent from unstructured human videos drawn from internet-scale datasets and learns robot control in massively-parallel simulation. The intent model predicts short-horizon intent (what should happen next in the scene) from the current observation in closed loop. An embodiment-specific sensorimotor policy converts this intent into robot actions. The intent interface is shared across controllers, so the same intent model can be applied to different embodiments, from our primary dexterous hand to a parallel-jaw gripper. We evaluate LUCID on five real-world manipulation tasks: stirring, wiping, and binning supervised by only internet video, with zero-shot transfer to novel scenes and object instances; and push-T and cable routing supervised by 1 hr each of self-collected smartphone video. Project page: https://lucid-robot.github.io/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.