TempoVLA: 속도 조절이 가능한 시각-언어-행동 정책 학습
TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
로봇 조작은 낮은 위험의 이동 단계와 빠른 실행을 요구하는 단계, 그리고 높은 위험의 접촉 단계와 느리고 정밀한 움직임을 요구하는 단계가 반복됩니다. 그러나 기존의 시각-언어-행동(VLA) 모델들은 훈련 데이터에서 얻은 단일 고정 속도만을 사용합니다. 모델 압축, KV 캐시 재사용 또는 강화 학습을 통해 VLA의 속도를 높이려는 이전 연구들은 정책을 단순히 한 가지 고정된 속도에서 다른 고정된 속도로 변경할 뿐이며, 감속 기능은 거의 탐구되지 않았습니다. 우리는 예측된 각 행동의 크기가 로봇의 움직임 속도를 결정한다는 것을 관찰했습니다. 이를 통해 로봇의 실행 속도를 명시적으로 제어할 수 있는 방법을 제시합니다. TempoVLA는 두 가지 결합된 구성 요소로 이루어진 단일 VLA 모델입니다. (1) 데이터 측면에서는 Variable-Speed Trajectory Augmentation (VSTA) 기술을 사용하여, 행동을 병합하거나 분리하여 훈련 데이터를 원하는 속도로 재조정하며, 움직임의 의미를 유지합니다. (2) 모델 측면에서는 conditioning 메커니즘을 통해 정책에 속도 정보를 제공합니다. 실험 결과 VSTA는 요청된 속도로 도달하면서 거의 무시할 만한 움직임 오류를 보입니다. 시뮬레이션 및 실제 작업 환경에서의 실험에서 TempoVLA는 양방향의 유연한 속도 제어를 달성하는 것으로 나타났습니다. 또한 VSTA는 데이터 활용도를 향상시켜 기본 성능을 더욱 높입니다. 나아가, 대규모 다중 모달 모델과의 협력을 통해 TempoVLA는 로봇이 낮은 위험 단계에서는 빠르게 움직이고, 높은 위험 단계에서는 느리게 움직이는 동적인 속도 제어를 구현할 수 있습니다.
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, precise motion. Yet existing Vision-Language-Action models (VLAs) only inherit a single fixed speed from training demonstrations. Prior efforts to accelerate VLAs through model compression, KV-cache reuse, or reinforcement learning only shift the policy from one fixed speed to another, and leave deceleration almost unexplored. We observe that the magnitude of each predicted action already governs how fast the robot moves, opening a direct route to controllable execution speed. We turn this observation into TempoVLA, a single VLA whose execution speed is controlled by an explicit condition. TempoVLA combines two coupled components. (1) A data-side Variable-Speed Trajectory Augmentation (VSTA) that re-times demonstration to any target speed by merging or splitting actions while preserving its motion semantics. (2) A model-side conditioning mechanism that feeds the speed to the policy. Statistics show that VSTA reaches the requested speed with negligible motion error. Experiments in simulation and on real-world tasks demonstrate that TempoVLA achieves flexible speed control in both directions, while VSTA additionally boosts the default $1\times$ performance via better data utilization. Furthermore, by cooperating with a large multimodal model, TempoVLA realizes dynamic speed control, accelerating through low-risk phases and decelerating for high-risk ones.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.