Skip to content

Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

It is shown that standard single-layer defenses each fail on their own and can even backfire, and called on the community to move beyond per-model alignment and toward composite safety mechanisms before multi-agent LLM systems are deployed at scale.

Abstract

Large language models (LLMs) are increasingly deployed in multi-agent systems where a principal agent decomposes tasks and delegates them to subordinate agents that may invoke external tools. Safety alignment, however, is still evaluated almost exclusively under a single-agent threat model, treating safety as a property of the individual LLM. We show that this assumption breaks down: \emph{individual safety alignment fails to transfer to multi-agent settings}. Two failure mechanisms emerge under delegation: \emph{responsibility diffusion} on the principal side and \emph{role-bias compliance} on the subordinate side, jointly converting language-level refusal into actionable harm. We refer to this phenomenon as \textit{delegated misalignment} and study it through a three-condition protocol across 6 frontier LLMs on 49 hazardous tasks. Delegation amplifies end-to-end harm substantially: DeepSeek-V3.2's full-execution rate rises from 30.6\% to 77.6\% once delegation is introduced, and the same model behaves very differently across roles (GPT-5: 22.5\% as a single agent vs.\ 61.2\% as a subordinate). Ablations further show that standard single-layer defenses each fail on their own and can even backfire. We call on the community to move beyond per-model alignment and toward composite safety mechanisms before multi-agent LLM systems are deployed at scale.

View source

Similar papers

Preprint Aug 2026

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

It is indicated that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.

Zeyuan Li, Lukas Petersson, Alessandro Acquisti et al. · 2 citations
#artificial intelligence Preprint Sep 2026

AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents

This work introduces AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety, and shows that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.

Tian-Zhuo Yang, Zi-Rui Mi, Yan-Tao Huang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

This work systematizes MAS security through an execution-centered analysis of 197 works, introducing an A-I-R framework that organizes attacks by adversary position, interaction interface, and resulting system-level risk, unifying otherwise fragmented attack mechanisms across MAS.

Rui Yang, Jun-Jie Xu, Zheng-Yu Liu et al. · 1 citation
#artificial intelligence Preprint Aug 2026

Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems

It is argued that agent security must be evaluated under an untrusted-model assumption: a correct system is one in which a fully prompt-injected agent still cannot exceed the authority explicitly delegated to it, and an authorization broker is implemented that closes the gap.

Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi · 0 citations
#artificial intelligence Preprint Sep 2026

The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State

Agent frameworks increasingly delegate work by forking sub-agents; a common default makes the child inherit the parent's full working context. We measure how the effect of inherited state changes with capability, where $C_m$ denotes clean fork-fresh accuracy. We compare 3 inheritance policies: Reset (fork fresh: base e...

Jun-Hao Hu, S. Ramachandran · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.