Skip to content

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

Sep 2026 · 0 citations · 53 references
Computer Science

TL;DR

The results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.

Abstract

Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.

View source

Similar papers

Preprint Aug 2026

Understanding Stage-Wise Utility-Risk Trade-offs in LLM Agent Memory

Control evaluations reveal that targeted poisoning risk varies across memory operations and motivate stage-aware evaluation and control of LLM-agent memory, showing that targeted poisoning risk varies across memory operations.

Chuan-Chao Zang, Zi-Jian Cao, Xiang-Tao Meng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by wh...

Ming-Xi Zou, Lang-Zhang Liang, Zhuo Wang et al. · 0 citations
Preprint Sep 2026

Inferring Hidden User Models from the Behavior of Personalized LLM Agents

Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more priva...

Hao-Yang Li, Ya-Xin Xiao, Qing-Qing Ye et al. · 0 citations
Review Open access Aug 2026

Memory Governance For AI Agents: Defending Against Cognitive State Traps

Memory Governance is introduced, a security-oriented framework that treats agent memory as a governed asset subject to continuous evaluation rather than passive storage that combines provenance tracking, weighted trust score with explicitly constrained weights, exponential confidence decay, and three-state quarantine c...

Ayush Jain · 0 citations
Preprint Aug 2026

When Agents Talk: Honeytokens under Shared Memory

During a 2026 cyber-capability evaluation, short-lived AI agents turned a shared package repository into persistent memory, passing exploit findings to later agents and rebuilding the channel after it was removed, raising a question for defensive deception: can a honeytoken be harmless to trusted agents without becomin...

Joshua S. Gans · 1 citation
Book Open access Aug 2026

EvoFEND: Dual Memory-Driven Self-Evolving Fake News Detection

Real-world fake news is inherently dynamic: evidence within an event accumulates and conflicts over time, while deceptive tactics shift across events. However, most prior work formulates detection as a static, one-shot classification problem over fixed snapshots. This mismatch ignores the lifecycle of news and leaves d...

Beizhe Hu, Qiang Sheng, Hao Mi et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.