2607.28391v1 Jul 30, 2026 cs.RO

TacWAM: 메커니즘 인지 촉각 예측을 위한 앵커 기반 월드 액션 모델

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

Yiding Ma
Yiding Ma
Citations: 23
h-index: 2
Lei Jin
Lei Jin
Citations: 61
h-index: 4
Xin Zhang
Xin Zhang
Citations: 85
h-index: 4
Chen Gao
Chen Gao
Citations: 75
h-index: 3
Wei Wu
Wei Wu
Citations: 88
h-index: 5
Yong Li
Yong Li
Citations: 38
h-index: 2

월드 액션 모델(WAM)은 미래 상태 예측과 로봇 동작 생성 기능을 결합하지만, 기존 접근 방식은 주로 시각적 정보에 의존합니다. 시각적 예측은 장면 구조와 객체 움직임을 파악하지만, 접촉이 많은 조작 과정에서 발생하는 힘, 변형, 전단력 및 미끄러짐에 대한 제한적인 정보를 제공합니다. 이러한 문제를 해결하기 위해 촉각적 미래 예측은 의미 있는 물리적 정보를 담아야 하며, 동시에 동작 생성에 있어 지나치게 중요한 정보가 되어서는 안 됩니다. 본 논문에서는 이러한 과제를 해결하기 위한 메커니즘 인지 촉각 WAM인 TacWAM을 제안합니다. 첫째, 공간적으로 정렬된 퓨전(SAF) 촉각 인코더를 사용하여 촉각 특징, 밀집된 힘 필드 및 변형 흐름을 공유되는 잠재 예측 공간으로 매핑하며, 양방향 힘과 토크 재구성을 통해 전반적인 접촉 정보를 유지합니다. 둘째, 촉각 히스토리 인코더는 시간적 맥락을 제공하여 미래 촉각 예측이 현재 촉각 관찰 범위를 넘어 힘과 변형의 변화를 반영하도록 합니다. 셋째, 앵커 기반 삼중 모드(AGT) 어텐션은 현재 시각 및 촉각 앵커, 미래 예측 토큰 및 동작 토큰을 분리하여, 미래 촉각 상태가 동작 생성에 직접적으로 영향을 미치지 않도록 하여 학습 과정을 감독합니다. 우리는 TacWAM을 섬세한 물체 잡기, 지속적인 표면 접촉, 동적 인-핸드 조작 등 4가지 실제 환경의 접촉이 많은 조작 작업에서 평가했습니다. TacWAM은 평균 성공률 75.0%를 달성하여 가장 강력한 기준 모델보다 37.5%p 더 높은 성능을 보였습니다. 단계별 분석 결과, 촉각 히스토리가 제거되거나 미래 예측 대상에 대한 접근이 제한될 경우 성능 저하가 발생함을 확인했습니다. 이러한 결과는 정보적인 촉각 표현과 배포 환경에 적합한 제약 조건을 결합하면 접촉 인지 동작 학습을 향상시킬 수 있음을 시사합니다.

Original Abstract

World Action Models (WAMs) combine future-state prediction with robot action generation, but existing approaches largely rely on visual futures. Visual prediction captures scene structure and object motion, yet provides limited supervision for force, deformation, shear, and slip during contact-rich manipulation. This creates two design requirements: tactile futures should carry meaningful physical information, and they should not become privileged cues for action generation. We present TacWAM, a mechanics-aware tactile WAM that addresses this challenge in three steps. First, a Spatially Aligned Fusion (SAF) Tactile Encoder maps tactile appearance, dense force fields, and deformation flow into a shared latent prediction space, with bilateral force and torque reconstruction preserving global contact information. Second, a tactile history encoder provides temporal context so future tactile prediction reflects how force and deformation change beyond the current tactile observation. Third, Anchor-Guided Tri-Modal (AGT) Attention separates current visual and tactile anchors, future prediction tokens, and action tokens, allowing future tactile states to supervise training without being directly read by the action branch. We evaluate TacWAM on four real-world contact-rich manipulation tasks covering fragile grasping, sustained surface contact, and dynamic in-hand manipulation. TacWAM achieves an average success rate of 75.0%, exceeding the strongest evaluated baseline by 37.5 percentage points. Staged ablations show consistent degradation when tactile history is removed and access to future prediction targets is relaxed. These results indicate that future tactile supervision can improve contact-aware action learning when combined with informative tactile representations and deployment-consistent information constraints.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!