Skip to content
Review Open access

Memory Governance For AI Agents: Defending Against Cognitive State Traps

Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

Memory Governance is introduced, a security-oriented framework that treats agent memory as a governed asset subject to continuous evaluation rather than passive storage that combines provenance tracking, weighted trust score with explicitly constrained weights, exponential confidence decay, and three-state quarantine containment to reduce the long-term influence of adversarial information.

Abstract

Long-horizon AI agents increasingly depend on persistent memory for planning and decision-making, creating an attack surface that existing defenses leave largely unaddressed. Prompt filtering and output validation protect individual interactions but offer no protection once adversarial information enters long-term storage. This paper introduces Memory Governance, a security-oriented framework that treats agent memory as a governed asset subject to continuous evaluation rather than passive storage. The framework combines provenance tracking, a weighted trust score with explicitly constrained weights, exponential confidence decay, and three-state quarantine containment to reduce the long-term influence of adversarial information. A Trust-Decay Memory Evaluation algorithm classifies each memory object as Trusted, Review Required, or Quarantined based on source reliability, validation history, and time-elapsed confidence. A discrete-time contamination propagation model, adapted from epidemiological dynamics, derives the condition μ > β under which governance controls drive contamination density to zero at steady state. Together, these mechanisms establish memory governance as an architectural security control rather than an interaction-level filter.

Read PDF

Similar papers

Review Sep 2026

MemSentry: A Framework for Detecting Persistent Memory Poisoning in Agentic AI

Agentic AI systems with persistent memory introduce a distinct attack surface known as memory poisoning, in which adversarially crafted content is stored in long-term memory and subsequently influences future agent behavior. Such attacks can suppress security alerts, facilitate privilege escalation, alter trust relatio...

Ayan Roy, K. Basu · 0 citations
Preprint Aug 2026

SafeCommit: Certifying When Memory-Grounded Agents May Safely Act

SafeCommit, a risk controlled layer between agent reasoning and external execution, is introduced, a calibrated set of plausible latent worlds from memory, observations, tool outputs, provenance, and policy constraints that permits a side effectful action only when a conformal action certificate shows that the action i...

M. Akewar, Ravi Ranjan · 2 citations
Preprint Sep 2026

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state long after their interactions. This creates a safety regime that cannot be characterized by evaluating model responses in isolation. Emergence World, is a continuously runn...

Deepak Akkil, Tamer Abuelsaad, Karthik Vikram et al. · 0 citations
Preprint Aug 2026

When Agents Talk: Honeytokens under Shared Memory

During a 2026 cyber-capability evaluation, short-lived AI agents turned a shared package repository into persistent memory, passing exploit findings to later agents and rebuilding the channel after it was removed, raising a question for defensive deception: can a honeytoken be harmless to trusted agents without becomin...

Joshua S. Gans · 1 citation
Review Jul 2026

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

Cyber-capable AI agents combine language models with tools, memory, and execution environments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environment...

A. B. Siddik · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.