Skip to content

When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents

Jul 2026 · arXiv.org · Vol abs/2607.05189 · 1 citation · 55 references
Computer Science

TL;DR

Results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.

Abstract

Persistent personal agents combine long-term memory with access to users'external environments, enabling personalized foreground assistance and proactive background execution. This integration also creates a new path to compromise: untrusted external content can be silently written into persistent memory and later reused as trusted state. We study this threat as stealth memory injection, in which a remote black-box adversary delivers a single email payload that must induce the agent to write poisoned memory, stay hidden in the agent's response to the user, and affect future behavior. We introduce WhisperBench, a 108-case benchmark spanning five risk categories and both fact and preference poisoning. Built on a real IMAP/SMTP workflow and an authentic email agent skill, it enables full-cycle evaluation of stealth memory injection attacks. To enable this black-box attack under single-email delivery and without runtime feedback, we propose MemGhost, a one-shot payload generation framework. MemGhost uses an environment proxy to emulate persistent-agent execution and an objective proxy to convert memory adoption and conversational stealth into dense rubric-based rewards, then trains the attacker policy with supervised fine-tuning and reinforcement learning. Across 56 held-out test cases, MemGhost achieves 87.5% end-to-end success on OpenClaw with GPT-5.4 and 71.4% on Claude Code SDK with Sonnet 4.6. It also transfers across personal-agent architectures (NanoClaw and Hermes Agent) and memory backends (filesystem and vector-based Mem0), and remains effective against input-level, model-level, and system-level defenses. These results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.

View source

Similar papers

Preprint Aug 2026

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

This work proposes InjecMEM, a novel memory injection attack paradigm that requires only a single interaction to steer later responses of related queries toward a pre-specified output and achieves reliable topic-conditioned retrieval and targeted generation.

Hanling Tian, Gengyu Zhang, Zeyang Sha et al. · 3 citations
#artificial intelligence Preprint Sep 2026

When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

Harness design has transformed the development of LLM-based agents by integrating memory, tool use, and runtime control. However, this design also introduces security and privacy risks because malicious instructions from external sources may be written into persistent memory and persist across sessions. To study this risk, we propose PMPA, a Persistent Memory Poisoning Attack against harness-based agents. PMPA embeds malicious instructions into benign external sources and induces the victim agent to write them into persistent memory without directly accessing to the agent framework. Once stored, the poisoned memory can be retrieved in later sessions, triggering additional malicious actions and causing privacy leakage. We evaluate PMPA on OpenClaw and Claude Code across different backbone LLMs, input modalities, and trigger scenarios. Across all settings, PMPA achieves average Injection Success Rate (ISR) and Cross-session Attack Success Rate (C-ASR) of 73.7%/ 55.5% on OpenClaw and 66.9%/ 81.7% on Claude Code, while preserving benign task performance on both systems. We further evaluate a targeted prompt-level defense and find that it can reduce memory injection in many settings, but provides limited protection once the persistent memory has been poisoned.

Shu-Huai Huang, Jing-Feng Zhang, Hong Jia · 0 citations
Preprint Aug 2026

Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw

This paper introduces MemCollusion, an automated red-teaming framework for constructing collusive memory poisoning attacks, and develops MoltLab, a controlled research reproduction of Moltbook, in which crafted platform content must first be observed and distilled into persistent memory before influencing the agent's behavior in a separate session.

Zheng Lin, Yuzhen Huang, Zhenxing Niu et al. · 0 citations
Preprint Aug 2026

MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents

This work proposes MAFIA, a query-only Memory Attack framework via probing and Factual Injection against Audit, tailored to this extended threat model, and introduces a placement strategy that ensures retrieval-competitive injection via memory probing, budget allocation, and scheduling.

Jiamin Chen, Yi-Sen Gao, Yan-Ping Li et al. · 0 citations
Jul 2026

Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents

The first extraction attack designed for this threat model, SPORE decouples the adversarial command from retrieval anchors by persisting the command in short-term memory and emitting semantically pure anchors in tool responses, demonstrating that memory isolation alone is insufficient and call for reexamining tool-side trust boundaries in agent architectures.

Xinyu Gao, Wenyu Chen, Xiangtao Meng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.