Skip to content
Preprint

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

Aug 2026 · 3 citations · 37 references
Computer Science

TL;DR

TrAct is proposed, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction, enabling more accurate world modeling and stronger robot generalization.

Abstract

Robot actions are inherently embodiment-specific and only weakly aligned with image-space visual changes, limiting their effectiveness as conditioning signals for robot world models. In contrast, visual tracks provide an embodiment-agnostic representation of how task-relevant points move through a scene, offering dense image-space guidance for accurate and spatially precise future video prediction. Building on this observation, we propose TrAct, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction. TrAct consists of three components: a Vision-Language-Action-and-Track model (VLAT) that jointly predicts candidate actions and corresponding visual tracks from the current observation and language instruction; a track-conditioned world model (TWM) that predicts future visual outcomes conditioned on the proposed tracks; and a vision-language reward model (VLAC) that scores the predicted outcomes. At inference time, VLAT generates candidate action-track pairs, TWM rolls out their visual consequences, and VLAC selects the track whose predicted outcome best satisfies the instruction; the action paired with the selected track is then executed by the robot. Experiments on the proposed LIBERO-INTEGRAL benchmark and real-world Franka manipulation show that TrAct improves success rates from 27% to 55% in simulation and from 49% to 76% on real-world tasks compared with the strong VLA baseline $\pi_{0.5}$. Furthermore, TWM consistently improves video prediction quality over the action-conditioned world model (AWM). These results demonstrate that visual tracks provide an effective shared interface between robot control and visual prediction, enabling more accurate world modeling and stronger robot generalization.

View source

Similar papers

Preprint Sep 2026

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data...

Jun-Yi Hu, Shuaihang Yuan, Jia-Zhao Liang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.

Shao-An Wang, Ao-Cheng Luo, Fei Huang et al. · 3 citations
#machine learning Preprint Aug 2026

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

LM-X is introduced, which organizes prediction across task, event, and motor scales without claiming anatomical correspondence and shows that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.

Jin Lou, Zhi Jing, Xu-Peng Wang et al. · 0 citations
Preprint Sep 2026

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction, establishes learning functionally structured intermediate...

Yu-Peng Zheng, Xiang Li, Song-En Gu et al. · 2 citations
Preprint Sep 2026

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO succes...

Jian-Man Lin, S. Shailesh, Zhong-Yi Luo et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.