2607.25895v1 Jul 28, 2026 cs.RO

HiFi-UMI: 고품질 UMI 데이터를 활용하여 실제 로봇에 적용 가능한 조작 정책 학습

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Di Zhang
Di Zhang
Citations: 22
h-index: 3
Weitao Zhou
Weitao Zhou
Citations: 250
h-index: 7
Simple AI Yuteng Wei
Simple AI Yuteng Wei
Citations: 0
h-index: 0
Jinming Ma
Jinming Ma
Citations: 18
h-index: 3
Jiawei Wang
Jiawei Wang
Citations: 0
h-index: 0
Yushen Zuo
Yushen Zuo
Citations: 74
h-index: 4
Ke Rui
Ke Rui
Citations: 0
h-index: 0
Minglei Li
Minglei Li
Citations: 0
h-index: 0
Jinhao Zhang
Jinhao Zhang
Citations: 8
h-index: 1
Zhi Pan
Zhi Pan
Citations: 0
h-index: 0
Xiang Wang
Xiang Wang
Citations: 0
h-index: 0
Haoran Jia
Haoran Jia
Citations: 0
h-index: 0
Hua Du
Hua Du
Citations: 11
h-index: 1
Zicheng Zeng
Zicheng Zeng
Citations: 17
h-index: 2
Jun Ma
Jun Ma
Citations: 0
h-index: 0
Guiyu Qin
Guiyu Qin
Citations: 0
h-index: 0
Xiaofei Li
Xiaofei Li
Citations: 0
h-index: 0

실제 로봇을 사용하여 학습하는 조작 정책은, 고품질이면서 확장 가능한 데이터의 부족으로 인해 어려움을 겪습니다. 실제 로봇을 이용한 원격 제어는 정확하지만 비용이 많이 들고, 로봇 없이 수집된 UMI 데이터는 쉽게 확장 가능하며, 현재는 주로 사전 학습에 사용되고 있으며, 실제 로봇 데이터를 사용하여 후처리합니다. 본 연구에서는 실제 로봇 데이터의 비중을 줄이는 대신, 로봇 없이 수집된 UMI 데이터의 품질을 향상시켜 후처리를 없앨 수 있는지 질문합니다. 우리는 HiFi-UMI라는 휴대 가능한 UMI 데이터 생성 시스템을 제안하며, 이 시스템은 궤적 정확도, 그리퍼 간 상대적인 자세, 동기화 및 시야를 최적화하도록 설계되었습니다. 구체적으로는 헤드 마운트 스테레오 관성 SLAM, 재구성되지 않은 상대적인 자세, 공유된 마이크로초 단위의 GPIO 트리거, 그리고 각 핸드당 2개의 광각 카메라 (약 200도 시야)를 사용합니다. HiFi-UMI는 외부 추적 시스템 없이도 3mm 수준의 작업 공간 내 엔드 이펙터 정확도를 달성합니다. 이러한 데이터를 사용하여 실제 로봇을 전혀 사용하지 않고 학습된 정책(post-training)이, 실제로 로봇에 바로 적용되어 다양한 시나리오에서 원격 제어와 유사한 성능을 보이는 것을 확인했습니다. 특히 StarVLA-QwenPI, OpenPI-pi_0.5 및 LingBot-VA라는 세 가지 모델에서 성공률 차이가 각각 -2.5%, +3.1% 및 -0.6%였습니다. 가장 뛰어난 성능의 정책은 정밀한 삽입 작업에서 85%의 성공률을 보였으며, 이는 원격 제어 기준이 평가 환경에서 수집되었고 HiFi-UMI 궤적이 전혀 사용되지 않았음에도 불구하고 달성된 결과입니다. 동일 데이터셋에서 4,000시간 동안 사전 학습을 진행하면, 10개의 새로운 작업에서 평균적으로 41%의 액션 오류 감소 효과를 얻었으며, StarVLA-QwenPI 모델에서는 실제 로봇 성공률이 추가로 18.1% 증가했습니다. 우리는 HiFi-UMI-2K라는 2,000시간 분량의 마이크로초 단위로 동기화된 초광각 시야 데이터셋을 공개합니다. 이 데이터는 자동으로 재구성되고 시뮬레이션 리플레이를 통해 검증되었으며, 로봇 학습 커뮤니티를 위한 대규모 고품질 자원입니다.

Original Abstract

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!