The results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
Abstract
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
Control evaluations reveal that targeted poisoning risk varies across memory operations and motivate stage-aware evaluation and control of LLM-agent memory, showing that targeted poisoning risk varies across memory operations.
Chuan-Chao Zang, Zi-Jian Cao, Xiang-Tao Meng et al.· 0 citations
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by wh...
Ming-Xi Zou, Lang-Zhang Liang, Zhuo Wang et al.· 0 citations
Recent personalized LLM agents increasingly transform information retained in memory into compressed or structured representations, which we call user models, to guide later decisions. When source wording is removed from the state reachable through the ordinary interface, these models are commonly treated as more priva...
Hao-Yang Li, Ya-Xin Xiao, Qing-Qing Ye et al.· 0 citations
Memory Governance is introduced, a security-oriented framework that treats agent memory as a governed asset subject to continuous evaluation rather than passive storage that combines provenance tracking, weighted trust score with explicitly constrained weights, exponential confidence decay, and three-state quarantine c...
Ayush Jain· International journal of com...· 0 citations
During a 2026 cyber-capability evaluation, short-lived AI agents turned a shared package repository into persistent memory, passing exploit findings to later agents and rebuilding the channel after it was removed, raising a question for defensive deception: can a honeytoken be harmless to trusted agents without becomin...
Real-world fake news is inherently dynamic: evidence within an event accumulates and conflicts over time, while deceptive tactics shift across events. However, most prior work formulates detection as a static, one-shot classification problem over fixed snapshots. This mismatch ignores the lifecycle of news and leaves d...
Beizhe Hu, Qiang Sheng, Hao Mi et al.· Proceedings of the 32nd ACM...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.