Skip to content
Review

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation

Jul 2026 · arXiv.org · Vol abs/2607.28226 · 3 citations · 183 references
Computer Science

TL;DR

This survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

Abstract

World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.

View source

Similar papers

Review Aug 2026

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

This work presents a trust-boundary-centric survey of foundation-model-powered embodied-agent security, and shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection.

Jiawei Liu, Jia-Cheng Guo, Tian Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to s...

Wen-Kai Huang, Si-Yuan Liang, Gaolei Li et al. · 0 citations
Review Jul 2026

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

Cyber-capable AI agents combine language models with tools, memory, and execution environments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environment...

A. B. Siddik · 0 citations
Jul 2026

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational controls to achieve security-relevant compromise. This paradigm remains necessary for AI-enabled systems, but it is no longer sufficient. In such systems, adversaries may in...

Mohammad Allahbakhsh, Mohammad Bahari, Moslem Attar-Raouf · 0 citations
#artificial intelligence Review Open access 2026

Adversarial Machine Learning: Security Risks and Defense Strategies in AI-Driven Applications

A detailed overview of the security risks associated with adversarial attacks is offered, including evasion attacks carried out at inference time, data poisoning that corrupts the training process, backdoor insertion that hides dormant triggers inside a model, and model inversion that leaks private information back out...

Harsh Verma · 1 citation
Open access Jul 2026

Invariant-Centered, Agent-Assisted Defense: A Security Architecture for Unbounded Attack Techniques and Unenumerable Attack Objectives

A defensive architecture for under unbounded attack techniques and unenumerable attack objectives, where the highest-value security investments are those that hold regardless of technique, and relocates residual reactivity to a single measurable point.

Ajayi Abisoye, N. Hussain, Abolaji Adebayo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.