2607.25397v1 Jul 28, 2026 cs.RO

분해 및 재구성: 시각-운동 정책과 원시 동작을 활용한 계획 수립

Decompose and Reorganize: Planning with Primitives and Visuomotor Policies Learned from Demonstrations

Yizhou Chen
Yizhou Chen
Citations: 44
h-index: 3
Hang Xu
Hang Xu
Citations: 13
h-index: 2
Dongjie Yu
Dongjie Yu
Citations: 76
h-index: 3
Yupu Lu
Yupu Lu
Citations: 119
h-index: 6
Tengye Xu
Tengye Xu
Citations: 11
h-index: 2
Zeqing Zhang
Zeqing Zhang
Citations: 95
h-index: 4
Wei Zhang
Wei Zhang
Citations: 12
h-index: 3
Yi Ren
Yi Ren
Citations: 48
h-index: 2
Ben M. Chen
Ben M. Chen
Citations: 186
h-index: 7
Jia Pan
Jia Pan
Citations: 40
h-index: 3

숙련된 로봇 조작 작업을 자동화하려면 고수준 추론 능력과 정밀한 실행 능력을 모두 갖춘 프레임워크가 필요합니다. 전통적인 작업 및 동작 계획(TAMP)은 상징적 계획에 뛰어나지만, 접촉이 많은 작업에서는 취약점을 보이는 경우가 많습니다. 반면, 시각 피드백을 활용하는 조작 작업에서 효과적인 모방 학습(IL)은 공간 일반화 능력과 다단계 작동 능력 부족이라는 한계를 가지고 있습니다. 이러한 상호 보완적인 강점과 약점을 결합하기 위해, 우리는 DR-LfD (Decomposed and Reorganized Skills Learned from Demonstrations)라는 프레임워크를 제안합니다. DR-LfD는 시각-운동 정책을 TAMP 기반의 의사 결정 시스템에 통합하여 원활하게 작동하도록 합니다. DR-LfD는 접촉 관계를 기반으로 인간의 동작 시연을 기본 동작으로 분해하고, 이를 시각-운동 정책 또는 객체 중심 원시 동작으로 재현합니다. 시각-운동 정책의 시작, 종료 및 제약 조건을 신중하게 모델링하고 구현하여 TAMP와 호환되도록 하며, 이를 통해 다양한 출처에서 학습된 동작들을 재구성할 수 있습니다. DR-LfD는 문제 해결 방식을 복잡한 동작 순서에 대한 지수적인 양의 시연 데이터가 필요했던 방식에서, 각 동작 유형별로 제한된 데이터를 사용하여 전체 시연 데이터 요구량을 줄이는 방식으로 전환합니다. 다양한 실제 환경 및 시뮬레이션 벤치마킹을 통해 DR-LfD는 여러 단계가 포함되고, 이전에 보지 못한 환경 조건과 물리적 제약이 존재하는 작업에서 뛰어난 성능을 보여줍니다. 프로젝트 웹사이트: https://dr-lfd.github.io/DR-LfD-website.

Original Abstract

Successfully automating dexterous, long-horizon robotic manipulation requires frameworks capable of both high-level reasoning and fine-grained execution. Traditional task and motion planning (TAMP), while excellent at symbolic planning, is often brittle in contact-rich operations. Simultaneously, imitation learning (IL), while effective in manipulation tasks with visual feedback, is limited by its low capability in spatial generalization and multi-stage operation. To reconcile their complementary strengths and limitations, we propose DR-LfD (Decomposed and Reorganized Skills Learned from Demonstrations), a framework that seamlessly integrates visuomotor policies into a TAMP-gated decision-making system. Based on contact relationships, DR-LfD decomposes human demonstrations into atomic skills, which are reproduced as visuomotor policies or object-centric primitives. The initiation, termination, and constraints of the visuomotor policies are carefully modeled and implemented in a TAMP-compatible form, enabling reorganization of skills learned from different sources. DR-LfD transforms the learning problem from one requiring exponential demonstration data over possible skill sequences to one whose demonstration burden scales with the number of distinct skill types, with limited data for each skill. Through comprehensive real-world and simulation benchmarking across diverse scenarios, we demonstrate the strong performance of DR-LfD on tasks involving multiple steps, unseen setups, and physical constraints. Project website: https://dr-lfd.github.io/DR-LfD-website.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!