2607.20345v1 Jul 22, 2026 cs.RO

연구실과 실제 매장 간 격차 해소: 소매용 휴머노이드 로봇을 위한 데이터 효율적인 사후 학습 및 경험 기반 학습 VLA 프레임워크

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

Tran Nguyen Le
Tran Nguyen Le
Citations: 6
h-index: 2
Roger Sala Sisó
Roger Sala Sisó
Citations: 0
h-index: 0
Tiago Silvério
Tiago Silvério
Citations: 55
h-index: 4
J. Sand
J. Sand
Citations: 1
h-index: 1

비전-언어-행동(VLA) 휴머노이드 로봇은 실행 오류, 데이터 분포 변화 및 환경 변동성을 극복해야 하며, 이는 벤치마크 성능과 실제 운영 간의 격차를 줄이는 데 있어 핵심적인 과제입니다. 본 논문에서는 Unitree G1-Edu 휴머노이드 로봇과 GR00T N1.6 기반 모델을 사용하여 슈퍼마켓 스낵 보충 작업을 수행하는 시스템 수준 접근 방식인 DEED (데이터 효율적인 사후 학습 및 경험 기반 학습)를 제안합니다. DEED는 세 가지 주요 구성 요소로 이루어져 있습니다: (1) 제어 빈도 정렬, 데이터 큐레이션, 작업 관련 시각적 강조 및 VLA 의존성 감소를 포함하는 데이터 효율적인 사후 학습 파이프라인; (2) 텍스트 기반 장점 접두사와 비전-언어 가치 함수를 통해 RECAP에서 영감을 받은 경험 기반 개선에 대한 실제 연구; (3) 분포 내/외부 동작을 분석하기 위한 잠재 공간 분석 도구. 우리의 결과는 연구실과 실제 매장 간의 격차를 해소하는 것이 주로 시스템 통합 문제이며, 특정 아키텍처 문제가 아니라는 것을 시사합니다. 신중한 데이터 설계 및 목표 지향적인 사후 학습은 단일 GPU만으로 초기 성능이 좋지 않은 정책을 유능한 실제 운영 시스템으로 변환할 수 있습니다.

Original Abstract

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!