τ: 미래 시각적 감독을 활용한 촉각 증강 시각-언어-행동 모델 학습
τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
시각-언어-행동(VLA) 모델에 촉각 센싱 기능을 통합하면 접촉이 빈번하게 발생하는 조작 작업에서 성능 향상을 기대할 수 있습니다. 왜냐하면 시각 정보만으로는 물리적 상호 작용에 대한 중요한 단서를 파악하기 어렵기 때문입니다. 그러나 제한된 작업별 데이터 하에서도 유용한 촉각 표현을 학습하고, 이를 사전 훈련된 VLA 모델에 효과적으로 적용하는 것은 여전히 어려운 과제입니다. 기존 방법들은 주로 순간적인 접촉 상태에 집중하거나, 6차원 힘(wrench) 시퀀스를 사용하여 시간적 상호 작용 역학을 모델링하지만, 고차원의 촉각 신호는 충분히 활용되지 못합니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 미래 시각적 감독을 통해 작업에 조건화된 공간-시간 촉각 표현을 학습하고, 이를 JEPA(Joint-Embedding Predictive Architecture)에서 영감을 받아 VLA 프레임워크인 τ를 제안합니다. 이 방법은 잠재 공간에서 작동하며 학습 과정에서만 사용되므로, 실제 적용 시 추가적인 오버헤드가 발생하지 않습니다. 또한, 본 논문에서는 네 가지 대표적인 접촉이 빈번한 조작 작업에 대한 동기화된 시각, 고유수용성 및 시각 기반 촉각 신호를 포함하는 데이터셋인 TacAura를 공개합니다. 실험 결과는 τ가 기존 모델보다 우수한 성능을 보이며, 새로운 객체와 환경에서도 일반화되어 향상된 조작 성능과 안정성을 제공함을 보여줍니다.
Incorporating tactile sensing into Vision-Language-Action (VLA) models holds promise for contact-rich manipulation, where visual observations alone often fail to capture critical cues about physical interactions. However, learning informative tactile representation while effectively adapting it to pretrained VLA models remains challenging under limited task-specific data. Existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present τ, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that τ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness. Project Page: https://cocacola-lab.github.io/tau-Page/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.