2608.04196v1 Aug 04, 2026 cs.RO

SiMDex: 로봇의 숙련된 조작을 위한 유사한 시점 영상 데이터 마이닝

SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Xiao Ma
Xiao Ma
Citations: 861
h-index: 13
Nie Lin
Nie Lin
Citations: 76
h-index: 3
Takehiko Ohkawa
Takehiko Ohkawa
Citations: 246
h-index: 8
Ruoshi Wen
Ruoshi Wen
Citations: 191
h-index: 8
Zhengming Zhu
Zhengming Zhu
Citations: 16
h-index: 2
Yiming Bao
Yiming Bao
Citations: 0
h-index: 0
Minjie Cai
Minjie Cai
Citations: 877
h-index: 14
Wei Xu
Wei Xu
Citations: 16
h-index: 2
Yoichi Sato
Yoichi Sato
Citations: 105
h-index: 4
Sijin Chen
Sijin Chen
Citations: 142
h-index: 5
Zhuohang Li
Zhuohang Li
Citations: 95
h-index: 4
Liqun Huang
Liqun Huang
Citations: 164
h-index: 6

최근 몇 년 동안, 로봇 조작을 위한 시점 기반 인간 영상을 활용하는 연구가 급증하고 있지만, 어떤 데이터가 실제로 숙련된 조작에 도움이 되는지에 대한 명확성은 여전히 부족합니다. 본 논문에서는 SiMDex라는 유사성 기반 데이터 마이닝 프레임워크를 제시하며, 이는 VLA(Visual-Language-Action) 후속 학습을 위한 인간 데이터 선택 문제를 추천 문제로 재구성하는 것입니다. SiMDex는 각 로봇 데모에 대해 3200만 개의 시점 기반 인간 샘플 풀에서 작업 관련 하위 집합을 추출하기 위해 세 단계의 (검색-순위 결정-재순위 결정) 파이프라인을 사용합니다. 이 프레임워크는 형태론적으로 독립적인 동작 공간에서 작동하며, VLA 아키텍처나 학습 과정에 변경 사항이 필요하지 않습니다. 강력한 기준 모델과 비교했을 때, SiMDex는 무작위로 샘플링된 인간 데이터와 동일한 양의 데이터를 사용하면서도 전체 성공률을 47.7%에서 61.1%로 향상시킵니다. 이는 선별적인 데이터 큐레이션이 무분별한 데이터 혼합보다 우수하다는 것을 보여줍니다. SiMDex는 전체 풀의 약 5%에 해당하는 149만 개의 마이닝된 샘플만을 사용합니다.

Original Abstract

Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity-based data mining framework that casts human data selection for VLA post-training in dexterous manipulation as a recommendation problem. For each robot demonstration, SiMDex employs a three-layer recall-ranking-re-ranking pipeline to extract task-relevant subsets from a pool of ~32M egocentric human samples, operating in a morphology-agnostic action space that requires no changes to VLA architecture or training. Against a strong baseline trained with an equal amount of randomly sampled human data, SiMDex uses only ~1.49M mined samples (<5% of the pool) yet improves the overall success rate from 47.7% to 61.1%, showing that selective curation outperforms indiscriminate data mixing.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!