Skip to content

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

Jul 2026 · arXiv.org · Vol abs/2607.22962 · 2 citations · 28 references
Computer Science

TL;DR

A write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments.

Abstract

LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for every subsequent step, a failure mode we call memory contamination. Existing memory management addresses retrieval and capacity but not write-time correctness; this admission problem cannot be solved by utility- or recency-based criteria, and uncontrolled contamination compounds across long trajectories. We propose ConsistencyGate, a write-time admission gate that, before committing a candidate fact m extracted from context c, queries the LLM K times for a soft support score and admits m only when the average exceeds a threshold. The mechanism is model-agnostic, requires no fine-tuning, and reduces to a single forward pass in a log-probability variant for latency-sensitive deployments. To measure the effect on natural data, we construct two real-conversation benchmarks (LoCoMo-Contam and MSC-Contam) by planting controlled single-detail corruptions in long-term conversations from LoCoMo and MSC, and complement them with a structured synthetic corpus (MemContam) that isolates a near-oracle upper bound. Across four LLM backbones, ConsistencyGate reduces contamination on every benchmark relative to a write-everything baseline, with the cost concentrated on facts that are stated only implicitly in the source context. We release all three benchmarks together with the gate implementation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

The Immutable Past: Formalizing State Mutability and Conflict Resolution in Mutable RAG

Retrieval-Augmented Generation (RAG) serves as the primary memory architecture for long-horizon autonomous agents. However, treating shared memory as an append-only stream introduces \textit{Semantic Shadowing}, a critical failure mode where conflicting historical observations accumulate and statistically dominate vali...

Hamed HaddadPajouh, Amir AmiriTabat · 0 citations
Preprint Aug 2026

MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance

LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks. Such memory is useful only when stored experience remains reliable across hundreds of interactions, but two failure modes break that assumption in pract...

Hao-Yu Wang, Guang-Yuan Dong, He Liang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Agent Memory Is a Surface for Endogenous Authorization Laundering

This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.

Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol · 3 citations · ⚡1
Preprint Aug 2026

MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents

This work proposes MAFIA, a query-only Memory Attack framework via probing and Factual Injection against Audit, tailored to this extended threat model, and introduces a placement strategy that ensures retrieval-competitive injection via memory probing, budget allocation, and scheduling.

Jiamin Chen, Yi-Sen Gao, Yan-Ping Li et al. · 0 citations
Preprint Aug 2026

SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents

We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents, unifying multi-channel retrieval under reciprocal-rank fusion, bi-temporal recall, multi-scope isolation, role-based access, verified erasure, and a hash-chained audit trail. A reliability spine governs the primary write pat...

V. Bhardwaj, Garima Singh, Arun Pratap Bhardwaj · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.