RoboTTT: 로봇 제어를 위한 컨텍스트 확장
RoboTTT: Context Scaling for Robot Policies
최근의 로봇 기반 모델은 주로 단일 단계 또는 짧은 시계열의 시각-운동 정보를 사용합니다. 본 연구에서는 Test-Time-Training Robot Policies (RoboTTT)라는 로봇 모델과 훈련 방법을 소개하며, 이는 기존 최고 성능 모델보다 훨씬 긴 8K 타임스텝 길이의 시각-운동 컨텍스트를 활용하면서도 추론 지연 시간을 늘리지 않습니다. 이러한 긴 컨텍스트 길이를 통해 새로운 로봇 기능을 구현할 수 있습니다: 인간 비디오 데모에서 즉석(one-shot) 모방, 실시간 정책 개선, 외부 간섭에 대한 강건성 확보, 그리고 다단계 및 장기 목표를 가진 작업에서의 성능 향상입니다. 또한, 사전 훈련 컨텍스트 길이가 증가함에 따라 폐루프 성능이 꾸준히 향상되는 현상을 처음으로 관찰했습니다. RoboTTT는 Vision-Language-Action 정책과 같은 로봇 기반 모델에 Test-Time Training을 통합하여 설계되었으며, 이 모델은 빠른 가중치로 구성된 순환 상태를 가지며, 훈련 및 추론 과정에서 경사 하강법을 통해 업데이트되는 파라미터를 사용하여 과거 정보를 가중치 공간에 압축하고 장기 컨텍스트 조건 설정에 필요한 문맥 정보를 검색합니다. 훈련 컨텍스트 길이를 확장하기 위해, 본 연구에서는 순차적 액션 강제와 절단된 시간 역전파를 결합한 방법을 사용했습니다. 어려운 실제 로봇 조작 작업에서 RoboTTT는 단일 단계 컨텍스트 기준 모델보다 전체 성능이 87% 향상되었으며, 기존 모델로는 완료할 수 없었던 5분 동안의 10단계 조립 작업을 완전히 성공적으로 수행했습니다. 8K 타임스텝으로 훈련된 RoboTTT는 1K 타임스텝으로 사전 훈련된 동일한 모델보다 62% 더 높은 성능을 보였으며, 이는 로봇 기반 모델에서 컨텍스트 길이를 새로운 성능 확장 요소로 제시합니다. 관련 영상 자료는 https://research.nvidia.com/labs/gear/robottt/ 에서 확인할 수 있습니다.
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.