Skip to content

RoboTTT: Context Scaling for Robot Policies

Jul 2026 · arXiv.org · Vol abs/2607.15275 · 9 citations
Computer Science

TL;DR

RoboTTT, a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency, unlocks new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks.

Abstract

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/

View source

Similar papers

Preprint Aug 2026

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.

Chunkai Yang, An-Dong Yang, Di Huang et al. · 0 citations
Preprint Aug 2026

PredVLA: Predictive Sensorimotor Modeling for Sub-Million-Parameter Robot Manipulation

A mechanism-by-mechanism transition to the recurrent behavior-cloning baseline shows that replacing the predictive pathway with direct observation input produces the largest single performance drop, accounting for approximately $70\% of the endpoint gap.

Hiroki Sawada, Shunichi Kasahara · 0 citations
Jul 2026

BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning

BWM is an action-conditioned world model that combines initial-environment guidance, dynamic visual history, and temporally aligned robot-action conditioning for stateful autoregressive prediction of future observations and is released as an open-source, low-cost, high-fidelity world simulator for robot manipulation.

Bwm Team · 1 citation
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations

Lightweight Adaptation of Pretrained Robot Manipulation Systems: Two Approaches

Two systematic attempts to improve large pretrained models with minimal or zero modification to their weights via reinforcement learning on a frozen OpenVLA-7B using binary task-success rewards on LIBERO-Goal reveal a common ceiling.

Adam Lalani, Chen Sun, Hui Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.