Skip to content

SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry

Aug 2026 · 3 citations · ⚡ 1 influential · 24 references
Computer Science

TL;DR

The results show that planning depends not only on whether task-relevant information is present, but also on whether it shapes the geometry consumed by the planner.

Abstract

Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself. Two prominent strategies for obtaining non-collapsed representations are to inherit a pretrained feature space, as in DINO-WM, and to learn an embedding end to end with anti-collapse regularization, as in LeWorldModel (LeWM) with SIGReg. These strategies show complementary strengths across tasks. Although task-relevant state is decodable from the full embeddings of both models, DINO-WM's leading principal components usually retain substantially more state information than LeWM's. Because Euclidean planning costs are dominated by high-variance directions, this difference affects how strongly state can influence candidate selection. We propose SCALE (State-CAlibrated Latent Embeddings) to give the end-to-end LeWM representation the favorable geometric property observed in DINO-WM. SCALE induces this property by correlating sampled pairwise latent distances with distances in a standardized task-relevant state space, without replacing LeWM's learned encoder. Across five tasks, three planning solvers, and five compute budgets, SCALE improves every task--solver average over LeWM. A latent-to-state regression control matches or exceeds SCALE's full-embedding decodability yet leaves latent--state distance alignment essentially unchanged and yields less consistent planning gains. SCALE adds a single lightweight training-time regularizer and no planning-time overhead. These results show that planning depends not only on whether task-relevant information is present, but also on whether it shapes the geometry consumed by the planner.

View source

Similar papers

#machine learning Preprint Sep 2026

Adaptive Latent Capacity for World Models

Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation, consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on avera...

Idan Achituve, Lior Dikstein, Idit Diamant et al. · 0 citations
Preprint Aug 2026

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

The Veracity-Influence-Sobriety score (VIScore), a metric that quantifies the reachability and capacity of a predictor given the encoded feature, and the hallucination of the searching-based planner, is proposed, showcasing the importance of these three aspects in planning success.

Hai-Yu Wu, Randall Balestriero, Morgan E. Levine · 3 citations
#artificial intelligence Preprint Aug 2026

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and it is argued that the anti-collapse pressure can instead come from the transition data itself.

Jack Boylan, Chris Hokamp · 6 citations
#artificial intelligence Preprint Sep 2026

LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models

This work introduces LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent and learned-query residual-context embeddings and shows that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled.

Lu-Zhe Huang, Lei Chu, Jing-Yi Liang et al. · 0 citations
Preprint Aug 2026

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

UniJEPA is presented, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space and shows that the same latent space supports controllable abstraction.

Lan-Ji An, Da-Wei Liu, Jin Li et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.