2608.05369v1 Aug 05, 2026 cs.RO

세계에서 손목으로: 작업 조건에 따른 미래 손목 모델링을 통한 정밀 로봇 조작

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Fushuo Huo
Fushuo Huo
Citations: 736
h-index: 14
Zhengyang Yan
Zhengyang Yan
Citations: 44
h-index: 3
Haosong Peng
Haosong Peng
Citations: 48
h-index: 5
Yalun Dai
Yalun Dai
Citations: 253
h-index: 9
Tianyu Qi
Tianyu Qi
Citations: 44
h-index: 3
Yuhao Pan
Yuhao Pan
Citations: 49
h-index: 4
Zhengsheng Zhang
Zhengsheng Zhang
Citations: 87
h-index: 6
Chujie Wang
Chujie Wang
Citations: 3
h-index: 1
Xiucheng Wang
Xiucheng Wang
Citations: 988
h-index: 15
Nan Cheng
Nan Cheng
Citations: 8
h-index: 2
Wenchao Xu
Wenchao Xu
Citations: 64
h-index: 5

비전-언어-행동(VLA) 모델은 종종 주 시점과 손목 시점의 관찰 데이터를 평행적인 시각적 입력으로 처리하며, 이는 로봇 조작에서 이들의 뚜렷한 역할을 간과하는 결과를 초래합니다. 그러나 정밀 조작은 전반적인 작업 맥락 하에서 손목 부위에서의 상호 작용이 어떻게 변화될 것인지 예측하는데 크게 의존합니다. 이러한 한계를 해결하기 위해, 본 연구에서는 작업 조건에 따른 미래 손목 모델링을 수행하는 VLA 모델인 World-to-Wrist VLA (W2-VLA)를 제안합니다. W2-VLA는 현재의 다중 시점 관찰 데이터와 작업 지시 사항을 기반으로, 비전-언어 모델과 손목 예측기 사이의 간결한 인터페이스 역할을 하는 잠재적 모델링 토큰 집합을 맥락화합니다. 이 인터페이스와 관찰된 손목 기록에 따라, 예측기는 미래 손목 잠재 값을 예측하고, 이는 행동 예측을 위한 미래 정보를 포함하는 맥락으로 변환됩니다. 또한, 조작 진행 상황, 물리적 변화 단서 및 손목 부위 관련 증거를 설명하는 구조화된 주석을 생성하는 합성 파이프라인인 W2-CoT를 소개합니다. 이러한 주석은 보조적인 감독 신호를 제공하여 작업 조건에 따른 잠재 인터페이스 형성을 돕습니다. LIBERO, RoboTwin 2.0 및 실제 로봇 조작 작업을 통해 수행된 실험 결과는 단일 팔 및 양팔 설정 모두에서 정밀하고 접촉 감응적인 조작 성능이 향상되었으며, 동시에 행동 생성 속도는 80 Hz 이상으로 유지됨을 보여줍니다.

Original Abstract

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!