2603.19201v3 Mar 19, 2026 cs.RO

OmniVTA: 시각-촉각 세계 모델링을 통한 접촉 기반 로봇 조작

OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation

Ruihai Wu
Ruihai Wu
Citations: 25
h-index: 3
Yupeng Zheng
Yupeng Zheng
Citations: 872
h-index: 16
Yuhang Zheng
Yuhang Zheng
Citations: 439
h-index: 9
Chen Gao
Chen Gao
Citations: 68
h-index: 3
Yilun Chen
Yilun Chen
Citations: 6,221
h-index: 30
Songen Gu
Songen Gu
Citations: 212
h-index: 6
Weize Li
Weize Li
Citations: 19
h-index: 2
Yujie Zang
Yujie Zang
Citations: 34
h-index: 3
Shuai Tian
Shuai Tian
Citations: 153
h-index: 4
Xiang Li
Xiang Li
Institute for AI Industry Research (AIR), Tsinghua University
Citations: 91
h-index: 3
Ce Hao
Ce Hao
Citations: 166
h-index: 6
Si Liu
Si Liu
Citations: 139
h-index: 4
Haoran Li
Haoran Li
Citations: 111
h-index: 6
Shuicheng Yan
Shuicheng Yan
Citations: 171
h-index: 6
Wen-Juan Ding
Wen-Juan Ding
Citations: 11
h-index: 1

접촉이 많은 조작 작업, 예를 들어 닦기 및 조립은 정확한 접촉력 인지, 마찰 변화, 그리고 상태 전환을 필요로 하지만, 이러한 정보는 시각만으로는 신뢰성 있게 추론하기 어렵습니다. 시각-촉각 조작에 대한 관심이 높아지고 있지만, 기존 데이터셋의 규모가 작고 작업 범위가 좁으며, 현재 방법들은 촉각 신호를 수동적인 관찰 자료로 취급하여 접촉 역학을 모델링하거나 명시적으로 폐루프 제어를 가능하게 하는 데 활용하지 못한다는 두 가지 주요 한계점이 존재합니다. 본 논문에서는 대규모 시각-촉각-행동 데이터셋인 extbf{OmniViTac}을 제시합니다. 이 데이터셋은 86개의 작업과 100개 이상의 객체에 대한 21,000개 이상의 경로를 포함하며, 물리 법칙에 기반한 여섯 가지 상호작용 패턴으로 구성되어 있습니다. 이러한 데이터셋을 바탕으로, 본 논문에서는 extbf{OmniVTA}라는 시각-촉각 조작 프레임워크를 제안합니다. OmniVTA는 네 가지 핵심 모듈로 구성됩니다: 자기 지도 촉각 인코더, 짧은 시간 범위의 접촉 변화를 예측하기 위한 이중 스트림 시각-촉각 세계 모델, 행동 생성에 사용되는 접촉 인식 융합 정책, 그리고 예측된 촉각 신호와 실제 측정된 촉각 신호 간의 차이를 폐루프 방식으로 보정하는 60Hz 반사 제어기. 모든 상호작용 범주에서 실제 로봇 실험을 통해 OmniVTA가 기존 방법보다 우수한 성능을 보이며, 새로운 객체 및 기하학적 구성에 대한 일반화 능력이 뛰어나다는 것을 확인했습니다. 이는 예측 기반 접촉 모델링과 고주파 촉각 피드백을 결합하여 접촉 기반 조작 작업을 수행하는 데 매우 효과적이라는 것을 입증합니다. 모든 데이터, 모델, 코드는 프로젝트 웹사이트(https://mrsecant.github.io/OmniVTA)에서 공개적으로 제공될 예정입니다.

Original Abstract

Contact-rich manipulation tasks, such as wiping and assembly, require accurate perception of contact forces, friction changes, and state transitions that cannot be reliably inferred from vision alone. Despite growing interest in visuo-tactile manipulation, progress is constrained by two persistent limitations: existing datasets are small in scale and narrow in task coverage, and current methods treat tactile signals as passive observations rather than using them to model contact dynamics or enable closed-loop control explicitly. In this paper, we present \textbf{OmniViTac}, a large-scale visuo-tactile-action dataset comprising $21{,}000+$ trajectories across $86$ tasks and $100+$ objects, organized into six physics-grounded interaction patterns. Building on this dataset, we propose \textbf{OmniVTA}, a world-model-based visuo-tactile manipulation framework that integrates four tightly coupled modules: a self-supervised tactile encoder, a two-stream visuo-tactile world model for predicting short-horizon contact evolution, a contact-aware fusion policy for action generation, and a 60Hz reflexive controller that corrects deviations between predicted and observed tactile signals in a closed loop. Real-robot experiments across all six interaction categories show that OmniVTA outperforms existing methods and generalizes well to unseen objects and geometric configurations, confirming the value of combining predictive contact modeling with high-frequency tactile feedback for contact-rich manipulation. All data, models, and code will be made publicly available on the project website at https://mrsecant.github.io/OmniVTA.

21 Citations
1 Influential
15 Altmetric
98.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!