LiLo-VLA: 연결된 객체 중심 정책을 활용한 구성적 장기 계획 조작
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
범용 로봇은 비정형 환경에서 여러 가지 기계적 구조 변화(예: 물체 부착 또는 분리)를 포함하는 작업을 수행하는 장기 계획 조작 능력을 갖춰야 합니다. 비전-언어-행동(VLA) 모델은 다양한 기본 기술을 학습할 수 있는 잠재력을 제공하지만, 이러한 기술들을 순차적으로 연결하는 조합적 복잡성으로 인해 어려움을 겪으며, 환경 변화에 민감하여 연쇄적인 오류가 발생하기 쉽습니다. 이러한 문제점을 해결하기 위해, 저희는 새로운 장기 계획 작업을 학습한 적이 없어도 0-샷 일반화가 가능한 모듈형 프레임워크인 LiLo-VLA(Linked Local VLA)를 제안합니다. 저희의 접근 방식은 이동과 상호작용을 분리합니다. Reaching 모듈은 전체적인 움직임을 처리하고, Interaction 모듈은 객체 중심 VLA를 사용하여 관심 객체를 개별적으로 처리하여 관련 없는 시각적 특징에 대한 강건성을 확보하고 공간적 구성에 대한 불변성을 제공합니다. 특히, 이러한 모듈성은 동적인 재계획 및 기술 재사용을 통해 견고한 오류 복구를 가능하게 하여, 엔드-투-엔드 방식에서 흔히 발생하는 연쇄적인 오류를 효과적으로 완화합니다. 저희는 두 가지 어려운 작업 세트인 LIBERO-Long++과 Ultra-Long으로 구성된 21개의 작업 시뮬레이션 벤치마크를 소개합니다. 이러한 시뮬레이션에서 LiLo-VLA는 평균 69%의 성공률을 달성하여, Pi0.5보다 41% 높고 OpenVLA-OFT보다 67% 높은 성능을 보였습니다. 또한, 8개의 장기 계획 작업을 대상으로 한 실제 환경 평가에서는 평균 85%의 성공률을 보였습니다. 프로젝트 페이지: https://yy-gx.github.io/LiLo-VLA/.
General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing them and are prone to cascading failures due to environmental sensitivity. To address these challenges, we propose LiLo-VLA (Linked Local VLA), a modular framework capable of zero-shot generalization to novel long-horizon tasks without ever being trained on them. Our approach decouples transport from interaction: a Reaching Module handles global motion, while an Interaction Module employs an object-centric VLA to process isolated objects of interest, ensuring robustness against irrelevant visual features and invariance to spatial configurations. Crucially, this modularity facilitates robust failure recovery through dynamic replanning and skill reuse, effectively mitigating the cascading errors common in end-to-end approaches. We introduce a 21-task simulation benchmark consisting of two challenging suites: LIBERO-Long++ and Ultra-Long. In these simulations, LiLo-VLA achieves a 69% average success rate, outperforming Pi0.5 by 41% and OpenVLA-OFT by 67%. Furthermore, real-world evaluations across 8 long-horizon tasks demonstrate an average success rate of 85%. Project page: https://yy-gx.github.io/LiLo-VLA/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.