2607.15275v1 Jul 16, 2026 cs.RO

RoboTTT: 로봇 제어를 위한 컨텍스트 확장

RoboTTT: Context Scaling for Robot Policies

Yunhao Ge
Yunhao Ge
Citations: 1,332
h-index: 11
Ruijie Zheng
Ruijie Zheng
Citations: 1,429
h-index: 9
LinxiJimFan
LinxiJimFan
Citations: 1,361
h-index: 10
Fengyuan Hu
Fengyuan Hu
Citations: 1,266
h-index: 8
Yuke Zhu
Yuke Zhu
Citations: 1,669
h-index: 13
Fei-Fei Li
Fei-Fei Li
Citations: 88
h-index: 3
Jimmy Wu
Jimmy Wu
Citations: 164
h-index: 2
Yevgen Chebo-tar
Yevgen Chebo-tar
Citations: 62
h-index: 2
Tianyuan Dai
Tianyuan Dai
Citations: 55
h-index: 5
Yunfan Jiang
Yunfan Jiang
Stanford University
Citations: 3,590
h-index: 11
Scott Reed
Scott Reed
Citations: 1,380
h-index: 8

최근의 로봇 기반 모델은 주로 단일 단계 또는 짧은 시계열의 시각-운동 정보를 사용합니다. 본 연구에서는 Test-Time-Training Robot Policies (RoboTTT)라는 로봇 모델과 훈련 방법을 소개하며, 이는 기존 최고 성능 모델보다 훨씬 긴 8K 타임스텝 길이의 시각-운동 컨텍스트를 활용하면서도 추론 지연 시간을 늘리지 않습니다. 이러한 긴 컨텍스트 길이를 통해 새로운 로봇 기능을 구현할 수 있습니다: 인간 비디오 데모에서 즉석(one-shot) 모방, 실시간 정책 개선, 외부 간섭에 대한 강건성 확보, 그리고 다단계 및 장기 목표를 가진 작업에서의 성능 향상입니다. 또한, 사전 훈련 컨텍스트 길이가 증가함에 따라 폐루프 성능이 꾸준히 향상되는 현상을 처음으로 관찰했습니다. RoboTTT는 Vision-Language-Action 정책과 같은 로봇 기반 모델에 Test-Time Training을 통합하여 설계되었으며, 이 모델은 빠른 가중치로 구성된 순환 상태를 가지며, 훈련 및 추론 과정에서 경사 하강법을 통해 업데이트되는 파라미터를 사용하여 과거 정보를 가중치 공간에 압축하고 장기 컨텍스트 조건 설정에 필요한 문맥 정보를 검색합니다. 훈련 컨텍스트 길이를 확장하기 위해, 본 연구에서는 순차적 액션 강제와 절단된 시간 역전파를 결합한 방법을 사용했습니다. 어려운 실제 로봇 조작 작업에서 RoboTTT는 단일 단계 컨텍스트 기준 모델보다 전체 성능이 87% 향상되었으며, 기존 모델로는 완료할 수 없었던 5분 동안의 10단계 조립 작업을 완전히 성공적으로 수행했습니다. 8K 타임스텝으로 훈련된 RoboTTT는 1K 타임스텝으로 사전 훈련된 동일한 모델보다 62% 더 높은 성능을 보였으며, 이는 로봇 기반 모델에서 컨텍스트 길이를 새로운 성능 확장 요소로 제시합니다. 관련 영상 자료는 https://research.nvidia.com/labs/gear/robottt/ 에서 확인할 수 있습니다.

Original Abstract

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/

1 Citations
0 Influential
6.5 Altmetric
33.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!