인사이트(InSight): 지향 가능한 VLA를 활용한 자율적인 기술 습득
InSight: Self-Guided Skill Acquisition via Steerable VLAs
비전-언어-행동(VLA) 모델은 데모 데이터를 통해 조작 기술을 학습할 수 있지만, 그 능력은 학습 데이터에 포함된 기술로 제한됩니다. 본 논문에서는 VLA 모델을 기본 동작 수준에서 제어 가능하도록 만들어 자율적인 기술 습득을 가능하게 하는 프레임워크인 인사이트(InSight)를 제시합니다 (예: "그리퍼를 볼로 이동", "위로 들어올리기", "병 기울이기"). 인사이트는 크게 두 가지 단계로 구성됩니다. (1) VLM 계획 분해 및 엔드-이펙터 자세 정보를 활용하여 데모 데이터를 레이블링된 기본 동작으로 나누는 자동화된 분할 파이프라인을 통해 VLA 기본 동작의 제어 가능성을 확보하고, (2) VLM 가이드 데이터 순환 시스템을 구축하여 새로운 작업을 수행하는 데 필요한 누락된 기본 동작을 식별하고, VLM에서 제안하는 저수준 제어를 사용하여 해당 기본 동작에 대한 시연을 자율적으로 수행하며, 성공적인 시연 결과를 자동으로 레이블링하고 저장한 후 VLA 학습 데이터 세트에 통합합니다. 우리는 인사이트를 시뮬레이션 환경 및 실제 조작 작업(블록 뒤집기, 서랍 닫기, 쓸기, 비틀기, 따르기 등)에서 평가했습니다. 이러한 작업은 대상 기술에 대한 어떠한 인간의 시연 없이 진행되었습니다. 학습된 기본 동작은 추가적인 인간의 시연 없이 새로운, 장기간의 작업을 수행하기 위해 조합될 수 있습니다. 우리의 연구 결과는 기본 동작의 제어 가능성이 VLA 정책에서의 지속적인 기술 습득을 위한 실질적인 기반이 될 수 있음을 보여줍니다. 프로젝트 웹사이트: https://insight-vla.github.io.
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.