A joint identifiability condition for controlled world models with Gaussian latent states with Gaussian latent states is presented, which consists of two coupled components: representation identifiability and transition identifiability, and it is proved that when this condition holds, minimizing the LeJEPA-style predictive objective can recover both latent states and controlled dynamics in the sense of orthogonal transformation.
Abstract
World model serves as a promising tool to infer environment dynamics under high-dimensional observations and candidate actions. Recently, LeCun's JEPA provides a compelling framework for learning such models in representation space. Its action-conditioned extension plays a central role in visual control and latent-space planning, but leaves a fundamental question: can it recover the controlled dynamics from nonlinear observations? This paper presents a joint identifiability condition for controlled world models with Gaussian latent states, which consists of two coupled components: (1) representation identifiability and (2) transition identifiability. The former depends on the spectral separation property while the latter is related to non-degenerate variation of conditional action. We prove that when this condition holds, minimizing the LeJEPA-style predictive objective can recover both latent states and controlled dynamics in the sense of orthogonal transformation. We further prove that the upper bound of transition prediction error is inversely proportional to the spectral separation margin. We also characterize an attainable amplification of counterfactual prediction error that scales inversely with the weakest conditional action-excitation margin. The theoretical predictions are empirically supported across four nonlinear observation settings.
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a trained latent actually contain, and what decides this? We answer with controlled interventions in POKEWORLD, an interactive environment whose visually identical objects hide mass, drag, and contact stiffness. A certificate-gated protocol first certifies each parameter as recoverable from raw observations, then measures whether it enters the latent, so a null result can be attributed to the objective rather than to the environment. The resulting identifiability map has two organizing mechanisms and one frontier. Inputs limit what can be known, while prediction targets decide what is retained. Stiffness enters the latent only when touch is forecast ($R^2=0.50$, compared with $-0.02$ when the same signal is merely fused into the input), and under single-step prediction a vision-only latent discards even perfectly visible object state. Drag marks the frontier. It carries a recoverability certificate of 0.89 yet plateaus near 0.13 under every deterministic prediction objective we test, while a supervised head on the same trunk reaches 0.45. Parameters whose readout is slow and ratio-type under the sensed coordinates fall outside what these objectives acquire. On RH20T, an input-target factorial across scaling curves reproduces both mechanisms across two robots and 4,258 episodes. Every arm missing information or prediction pressure stays flat over a fivefold data range, and only the full multimodal objective forecasts force beyond a persistence baseline, with held-out gains that grow with scale. Objective structure determines which physical parameters a latent acquires, and additional data improves only the parameters it already acquires.
It is argued that the outstanding obstacle to deploying world models in systems that cannot fail -- power, thermal, process control -- is not predictive fidelity but verifiability, and a research agenda for physics-grounded, verifiable world models is outlined that unifies the two lineages.
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.
Haodong Yan, Jiaguang Zhu, Ming-Ming Jia et al.· 2 citations
It is proved that the planner's suboptimality is bounded by twice this discrepancy between the predicted and the true plan-cost at the plan the planner commits to, whereas the data-averaged prediction error neither bounds nor tracks it.
Hanzhe You, Yonggang Zhang, Maohao Ran et al.· 2 citations
Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.
Guo An, Zijing Wu, Honghua Dong et al.· 0 citations
State space models (SSMs) have emerged as a compelling alternative to transformers for visual recognition, offering linear computational complexity while maintaining competitive accuracy. However, the lack of interpretability tools designed specifically for the recurrent dynamics of these models remains a significant gap: existing saliency methods originate from convolutional or attention-based architectures and do not account for the sequential state propagation that governs information flow in SSMs. We introduce controllability analysis, a framework grounded in control theory that quantifies the structural influence of each input position on the internal state dynamics of vision SSMs. We derive two complementary indices: a Jacobian-based measure that tracks sensitivity propagation through the full sequence via an efficient $O(LN)$ backward recursion, and a Gramian-based measure that admits a closed-form, fully parallelizable solution for diagonal state matrices. Unlike gradient-based attribution methods that produce class-specific explanations, controllability indices measure intrinsic model properties that are invariant to the target class. Experiments on seven image classification benchmarks spanning diverse visual domains (microscopy, optical coherence tomography (OCT), dermatoscopy, radiography, natural images, apparel, and remote sensing) demonstrate that controllability analysis produces perfectly class-agnostic explanations (cross-class correlation of 1.0 compared to -0.45 to 0.18 for Grad-CAM across datasets) and reveals layerwise information flow patterns that vary across imaging domains. On the dataset where the model is best trained (BloodMNIST, 98% test accuracy), the proposed Jacobian index also outperforms Grad-CAM on standard insertion/deletion faithfulness (Jacobian $0.63 {\,}\pm {\,}0.02$ versus Grad-CAM $0.47 {\,}\pm {\,}0.03$ , mean ± standard deviation across three training seeds, with nonoverlapping per-seed 95% bootstrap CIs); on the lower accuracy datasets Grad-CAM remains stronger on this class-specific metric, a gap we discuss as a diagnostic of incomplete alignment between the model's internal dynamics and the classification task rather than a failure of the structural framework. A complementary internal-attention IoU metric, measuring alignment between each method's saliency and the model's own L2-magnitude attention pattern, shows controllability outperforming Grad-CAM on six of the seven datasets. Directional decomposition of the controllability profiles further shows that influence progressively sharpens with depth, transitioning from diffuse responses in early layers to sparse, high-magnitude peaks in deep layers that align with semantically meaningful structures. The proposed framework provides a principled foundation for understanding, debugging, and improving vision SSMs across application domains.
M. Mabrok, Yalda Zafari· IEEE Transactions on Neural...· 1 citation