Skip to content

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

Jul 2026 · arXiv.org · Vol abs/2607.06807 · 0 citations · 73 references
Computer Science

TL;DR

AcMAS is proposed, an activation-based framework for malicious-behavior detection in MAS that significantly outperforms graph-based baselines against stealthy attacks, with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale.

Abstract

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks and explicit graph-based modeling of the MAS topology and agent-level interactions. In practice, real-world attacks are becoming more semantically stealthy, while MAS execution is typically asynchronous without the temporal alignment assumed by graph-based propagation models. To address these limitations, we propose AcMAS, an activation-based framework for malicious-behavior detection in MAS. By analyzing internal reasoning states in the activation space of local agents, AcMAS detects even stealthy attacks in a synchronization-robust fashion, without relying on explicit interaction graphs. Moreover, our activation analysis provides critical signals to guide AcMAS in restoring the functionality of compromised agents, rather than the disruptive agent isolation commonly used by the state-of-the-art methods. Comprehensive evaluation demonstrates that AcMAS significantly outperforms graph-based baselines against stealthy attacks, by +0.22 F1 in synchronous settings (0.94 vs. 0.72) and by +0.55 F1 in asynchronous settings (0.93 vs. 0.38), with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

This work systematizes MAS security through an execution-centered analysis of 197 works, introducing an A-I-R framework that organizes attacks by adversary position, interaction interface, and resulting system-level risk, unifying otherwise fragmented attack mechanisms across MAS.

Rui Yang, Jun-Jie Xu, Zhengyu Liu et al. · 1 citation
Jul 2026

SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems

SafeFlow is proposed, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem and reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm succ...

Haowen Dai, Zonghao Ying, Wenfeng Li et al. · 0 citations
Jul 2026

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

This work builds a working instance on a hierarchical multi-agent system, runs it under benign and attacked conditions across five language models and two task domains, and measures how much of that warning rests on removable surface cues of the attack rather than on its distributed structure.

D. Arias, Dev Prashant Mistry, Ren Wang et al. · 0 citations
Preprint Aug 2026

SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains

This work introduces SynChain, a self-synthesized attack paradigm utilizing persistence-aware directed supervised fine-tuning to induce agents to create poisoned yet benign-looking artifacts, proving that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.

Fuyao Zhang, Jiaming Zhang, Che Wang et al. · 0 citations
Preprint Jul 2026

From Monoliths to Swarms: A Study of Attack Surface Evolution in the Transition to Multi-Agent Web Systems

This work proposes a taxonomy to categorize attack vectors specific to web-based MAS, accounting for vulnerabilities introduced or amplified by the involvement of multiple agents, and presents a test-bed WebMASLab to analyze web agent security against a fully external, web-only adversary.

Yashaswi Malla, Sandra Siby · 0 citations
Preprint Aug 2026

What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions

LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructions during LLM inference, thereby subverting the agent's decision. Exi...

Yichao Gao, Yumo Zhang, Yunhao Yao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.