2607.04591v1 Jul 06, 2026 cs.RO

시각-언어-행동 학습을 위한 단순-복잡 구조화된 시범 제시 방법

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

Xin Qiu
Xin Qiu
Citations: 51
h-index: 4
Yi Yu
Yi Yu
Citations: 14
h-index: 3

시각-언어-행동(VLA) 모델은 시각적 인식, 언어 이해 및 로봇 액션 생성 기능을 통합하여 로봇 조작 분야에서 뛰어난 성능을 보여주었습니다. 기존 연구는 주로 모델 아키텍처, 학습 전략 및 데이터셋 규모 개선에 초점을 맞추었지만, 시범 자료 수집 및 구성 방식에는 상대적으로 덜 주의를 기울였습니다. 본 논문에서는 시범 자료 구성이 정책 학습 효율성, 학습 안정성 및 정책 일반화에 직접적인 영향을 미치는 중요한 요소임에도 불구하고 간과되어 왔음을 지적합니다. 이러한 문제를 해결하기 위해, 이중 팔 로봇 플랫폼을 사용하여 VLA 학습을 위한 단순-복잡 구조화된 시범 수집 전략을 제안합니다. 본 연구에서는 다음 세 가지 기본 원칙에 따라 데이터를 체계적으로 구성합니다: (i) 복잡한 조작 작업을 점진적으로 학습 가능한 하위 기술로 분해, (ii) 불필요한 변동성을 줄이기 위해 상호 작용 환경 표준화, (iii) 작업의 복잡도가 점진적으로 증가하는 순서대로 시범 자료 구성. 이러한 구조화된 설계는 VLA 모델이 먼저 기본적인 조작 기술을 습득한 후 더욱 복잡한 작업 조합을 학습하도록 하여, 장기적인 조작 작업을 보다 효과적으로 학습할 수 있도록 합니다. 제안하는 전략은 블록 잡기 및 정렬, 그리고 타월 접기와 같은 두 가지 대표적인 로봇 조작 작업에 대해 평가되었습니다. 실험 결과는 전체 작업 완료율과 학습 안정성 측면에서 기존의 엔드-투-엔드 완전 작업 경로를 직접 수집하는 방법보다 일관되게 개선된 성능을 보여주었습니다. 이러한 결과는 시범 자료 구성이 VLA 학습에서 간과되어 왔지만 중요한 요소임을 강조하며, 효율적인 기술 습득, 확장 가능한 데이터셋 구축 및 장기적인 로봇 조작에 대한 실질적인 통찰력을 제공합니다.

Original Abstract

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, and dataset scale, while little attention has been paid to how demonstrations are collected and organized. We identify demonstration organization as a fundamental yet overlooked aspect of imitation learning, as it directly affects policy learning efficiency, training stability, and policy generalization. To address this gap, we propose a simple-to-complex structured demonstration collection strategy for VLA learning using a dual-arm robotic platform. Our approach systematically organizes data through three general principles: (i) decomposing complex manipulation tasks into progressively learnable sub-skills, (ii) standardizing the interaction environment to reduce unnecessary variability, and (iii) organizing demonstrations according to progressively increasing task complexity. This structured design enables VLA models to first acquire fundamental manipulation skills before learning increasingly complex task compositions, facilitating more effective learning of long-horizon manipulation tasks. We evaluate the proposed strategy on two representative robotic manipulation tasks: block grasping and sorting, and towel folding. Experimental results show consistent improvements in task success rate and training stability compared with the baseline method of directly collecting end-to-end complete task trajectories. These findings highlight demonstration organization as a previously underexplored but important factor in VLA learning and provide practical insights into efficient skill acquisition, scalable dataset construction, and long-horizon robotic manipulation.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!