Delegated Misalignment: How Multi-Agent Structures Amplify LLM Safety Risks
It is shown that standard single-layer defenses each fail on their own and can even backfire, and called on the community to move beyond per-model alignment and toward composite safety mechanisms before multi-agent LLM systems are deployed at scale.
Zong-Hao Ying, Jia-Qi Yan, Hui-Ze Luo et al.
· 0 citations