It is shown that standard single-layer defenses each fail on their own and can even backfire, and called on the community to move beyond per-model alignment and toward composite safety mechanisms before multi-agent LLM systems are deployed at scale.
Abstract
Large language models (LLMs) are increasingly deployed in multi-agent systems where a principal agent decomposes tasks and delegates them to subordinate agents that may invoke external tools. Safety alignment, however, is still evaluated almost exclusively under a single-agent threat model, treating safety as a property of the individual LLM. We show that this assumption breaks down: \emph{individual safety alignment fails to transfer to multi-agent settings}. Two failure mechanisms emerge under delegation: \emph{responsibility diffusion} on the principal side and \emph{role-bias compliance} on the subordinate side, jointly converting language-level refusal into actionable harm. We refer to this phenomenon as \textit{delegated misalignment} and study it through a three-condition protocol across 6 frontier LLMs on 49 hazardous tasks. Delegation amplifies end-to-end harm substantially: DeepSeek-V3.2's full-execution rate rises from 30.6\% to 77.6\% once delegation is introduced, and the same model behaves very differently across roles (GPT-5: 22.5\% as a single agent vs.\ 61.2\% as a subordinate). Ablations further show that standard single-layer defenses each fail on their own and can even backfire. We call on the community to move beyond per-model alignment and toward composite safety mechanisms before multi-agent LLM systems are deployed at scale.
It is indicated that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.
Zeyuan Li, Lukas Petersson, Alessandro Acquisti et al.· 2 citations
This work introduces AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety, and shows that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Tian-Zhuo Yang, Zi-Rui Mi, Yan-Tao Huang et al.· 1 citation
This work systematizes MAS security through an execution-centered analysis of 197 works, introducing an A-I-R framework that organizes attacks by adversary position, interaction interface, and resulting system-level risk, unifying otherwise fragmented attack mechanisms across MAS.
Rui Yang, Jun-Jie Xu, Zheng-Yu Liu et al.· 1 citation
It is argued that agent security must be evaluated under an untrusted-model assumption: a correct system is one in which a fully prompt-injected agent still cannot exceed the authority explicitly delegated to it, and an authorization broker is implemented that closes the gap.
Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi· 0 citations
Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base e...
Jun-Hao Hu, S. Ramachandran· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.