세계에서 손목으로: 작업 조건에 따른 미래 손목 모델링을 통한 정밀 로봇 조작
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
비전-언어-행동(VLA) 모델은 종종 주 시점과 손목 시점의 관찰 데이터를 평행적인 시각적 입력으로 처리하며, 이는 로봇 조작에서 이들의 뚜렷한 역할을 간과하는 결과를 초래합니다. 그러나 정밀 조작은 전반적인 작업 맥락 하에서 손목 부위에서의 상호 작용이 어떻게 변화될 것인지 예측하는데 크게 의존합니다. 이러한 한계를 해결하기 위해, 본 연구에서는 작업 조건에 따른 미래 손목 모델링을 수행하는 VLA 모델인 World-to-Wrist VLA (W2-VLA)를 제안합니다. W2-VLA는 현재의 다중 시점 관찰 데이터와 작업 지시 사항을 기반으로, 비전-언어 모델과 손목 예측기 사이의 간결한 인터페이스 역할을 하는 잠재적 모델링 토큰 집합을 맥락화합니다. 이 인터페이스와 관찰된 손목 기록에 따라, 예측기는 미래 손목 잠재 값을 예측하고, 이는 행동 예측을 위한 미래 정보를 포함하는 맥락으로 변환됩니다. 또한, 조작 진행 상황, 물리적 변화 단서 및 손목 부위 관련 증거를 설명하는 구조화된 주석을 생성하는 합성 파이프라인인 W2-CoT를 소개합니다. 이러한 주석은 보조적인 감독 신호를 제공하여 작업 조건에 따른 잠재 인터페이스 형성을 돕습니다. LIBERO, RoboTwin 2.0 및 실제 로봇 조작 작업을 통해 수행된 실험 결과는 단일 팔 및 양팔 설정 모두에서 정밀하고 접촉 감응적인 조작 성능이 향상되었으며, 동시에 행동 생성 속도는 80 Hz 이상으로 유지됨을 보여줍니다.
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.