Aug 2026· 3 citations· ⚡ 1 influential· 24 references
Computer Science
TL;DR
The results show that planning depends not only on whether task-relevant information is present, but also on whether it shapes the geometry consumed by the planner.
Abstract
Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself. Two prominent strategies for obtaining non-collapsed representations are to inherit a pretrained feature space, as in DINO-WM, and to learn an embedding end to end with anti-collapse regularization, as in LeWorldModel (LeWM) with SIGReg. These strategies show complementary strengths across tasks. Although task-relevant state is decodable from the full embeddings of both models, DINO-WM's leading principal components usually retain substantially more state information than LeWM's. Because Euclidean planning costs are dominated by high-variance directions, this difference affects how strongly state can influence candidate selection. We propose SCALE (State-CAlibrated Latent Embeddings) to give the end-to-end LeWM representation the favorable geometric property observed in DINO-WM. SCALE induces this property by correlating sampled pairwise latent distances with distances in a standardized task-relevant state space, without replacing LeWM's learned encoder. Across five tasks, three planning solvers, and five compute budgets, SCALE improves every task--solver average over LeWM. A latent-to-state regression control matches or exceeds SCALE's full-embedding decodability yet leaves latent--state distance alignment essentially unchanged and yields less consistent planning gains. SCALE adds a single lightweight training-time regularizer and no planning-time overhead. These results show that planning depends not only on whether task-relevant information is present, but also on whether it shapes the geometry consumed by the planner.
Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation, consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on avera...
Idan Achituve, Lior Dikstein, Idit Diamant et al.· 0 citations
The Veracity-Influence-Sobriety score (VIScore), a metric that quantifies the reachability and capacity of a predictor given the encoded feature, and the hallucination of the searching-based planner, is proposed, showcasing the importance of these three aspects in planning success.
Hai-Yu Wu, Randall Balestriero, Morgan E. Levine· 3 citations
Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and it is argued that the anti-collapse pressure can instead come from the transition data itself.
This work introduces LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent and learned-query residual-context embeddings and shows that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled.
Lu-Zhe Huang, Lei Chu, Jing-Yi Liang et al.· 0 citations
UniJEPA is presented, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space and shows that the same latent space supports controllable abstraction.
Findings support the use of the Fourier auxiliary head method to improve both overall success rate and data efficiency, while avoiding representation laziness in latent world models.
Peng Zhu, Salvatore Penachio, Kaustav Mukherjee et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026