Skip to content

Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

Jul 2026 · arXiv.org · Vol abs/2607.29482 · 0 citations · 30 references
Computer Science

TL;DR

Temporal Policy is introduced, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem and bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control.

Abstract

By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learning from Demonstration (LfD), but often suffer from high inference cost. This paper introduces Temporal Policy, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem. By initializing the generative flow at the robot's recent history, we explicitly couple past states to future action sequences. This data-dependent coupling reduces transport cost and produces straight vector fields. We validate Temporal Policy across visuomotor simulation benchmarks and on a physical Barrett WAM 2x 7DoF teleoperation platform. Our approach reduces transport costs by nearly an order of magnitude compared to noise-initialized baselines, achieving a 19.1 ms inference latency on a single NVIDIA RTX 4080. Crucially, these geometric and computational efficiencies are achieved while matching the success rates of state-of-the-art baselines. This simplified transport geometry bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control. The code is publicly available at https://github.com/dmiller12/TemporalPolicy.

View source

Similar papers

Preprint Aug 2026

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.

Chunkai Yang, An-Dong Yang, Di Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control

Learning visual policies for locomotion and manipulation requires coordinating contact with the environment and can incur substantial computation and GPU memory costs. First-order policy gradients (FoPG) reduce training cost through differentiable simulation, but local optimization can converge to unintended contact pa...

Yi-Lang Liu, Hao-Xiang You, Qian Wang et al. · 0 citations

Lightweight Adaptation of Pretrained Robot Manipulation Systems: Two Approaches

Two systematic attempts to improve large pretrained models with minimal or zero modification to their weights via reinforcement learning on a frozen OpenVLA-7B using binary task-success rewards on LIBERO-Goal reveal a common ceiling.

Adam Lalani, Chen Sun, Hui Wang · 0 citations
Preprint Aug 2026

CoDrift: Compositional Drifting for Offline Reinforcement Learning

This work proposes CoDrift, a compositional framework for one-step generative policy learning that combines three objective-level fields into a unified policy field that compares favorably with state-of-the-art methods and achieves the best average rank in both settings.

Xiewei Ni, Ruo-Feng Mei, Xiang-Yu Xu · 0 citations
#artificial intelligence Preprint Sep 2026

LePlanner: An Iterative Amortized Controller For World Models

World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet planning in these latent spaces typically relies on one of two costly approaches. Search-based planners such as CEM, MPPI, and iCEM optimize action sequences through many pr...

Saksham Bansal, Om Naphade, Chayan Aggarwal et al. · 0 citations
#machine learning Preprint Sep 2026

VGFM: Expressive Robot Policies via Dense Value Guidance in Flow Matching

Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-ste...

Prajwal Koirala, Mark E. Campbell · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.