Cyber-capable AI agents combine language models with tools, memory, and execution environments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthesizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command-and-control, and the speed of automated action. We use two separate preliminary incident records: the reported July 2026 Hugging Face/OpenAI evaluation breach and Anthropic's subsequent three-incident evaluation review. A comparative evidence protocol distinguishes record-specific factual claims from the shared systems lesson: the evaluation environment is itself part of the security boundary. Across the taxonomy and records, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.
This paper organizes the area into a structured taxonomy along five axes: the de-tection capability targeted, the analysis paradigm employed, the agent archi-tecture, the degree of autonomy, and the evaluation methodology.
Andi Xia· Poster Volume 0008 The 2026...· 0 citations
SafeClawArena is developed, a benchmark of 406 adversarial tasks executed in containerized replicas of real agent platforms with canary-marked credentials and evaluated via automated taint tracking across nine output channels, exposing the inadequacy of current defenses and suggesting directions for future hardening.
It is shown that the more consequential risks lie one layer down, in the protocol between agents and commerce services, and a platform-agnostic defense that drives the structural attack-success rate to zero for four of the five structural classes.
This work presents a trust-boundary-centric survey of foundation-model-powered embodied-agent security, and shows that attack research is concentrated on multimodal perception and action interfaces, while defenses are especially concentrated on action-level and runtime protection.
Jiawei Liu, Jiacheng Guo, Tianwei Zhang et al.· 0 citations
A four-layer taxonomy mapping 13 vulnerability types across perception, brain, action, and interaction layers is contributed, and seven open problems centered on containment are identified.
Md. Jafrin Hossain, Mohammad Arif Hossain, Nirwan Ansari· 0 citations
Open Security Benchmark is presented, a framework that benchmarks agentic AI on security posture management work and surfaces a curated enterprise environment that evaluates posture investigation across two modalities: text-to-SQL over a relational snapshot and each vendor's native API over a served instance of the same environment.
Gal Engelberg, Michael Arenzon, Leon Goldberg· 0 citations