Skip to content
Preprint

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

Aug 2026 · 1 citation · 19 references
Computer Science

Abstract

AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing away what made it checkable, who said it, how sure they were, when it held. We call that failure factwashing, and release factwash, an open-source write-time gate that catches it deterministically, with named flags and evidence rather than an LLM judge. Building it answers a practical question: when does a cheap check suffice, and when do you need a model? What decides is whether the property has a bounded surface-cue inventory. Explicit negation cues are close to enumerable, so a word list finishes and transfers, reaching 0.91 F1 on untuned text. Hedging and attribution have open-ended realizations, so vocabulary plateaus near half recall, and a one-question LLM witness recovers +17 and +15 points of cue-detection recall at equal precision. Deployed, that witness may only lower a verdict, so it buys precision rather than coverage. We measure cue detection on 105,596 independently annotated sentences. A blind-labelled corpus of memory writes then locates the failure: 55% of bad writes in conversational hearsay, 7% in business email (p<0.001), so the first deployment question is not which detector to use but whether the failure occurs at all. On unmodified mem0 2.0.7, the gate flags 5 of 8 hedged-hearsay writes.

View source

Similar papers

Preprint Aug 2026

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model...

Alex Kwon · 0 citations
Preprint Aug 2026

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

The resulting divergence catastrophic remembering is named: the inverse of catastrophic forgetting around which continual learning is organized, the inverse of catastrophic forgetting around which continual learning is organized.

Kushal Chakrabarti · 2 citations
Review Jul 2026

Can an AI Assistant Really Forget? Auditable Deletion from Addressable Memory

Certifying that a deletion did what it declared does not certify that the record left no trace: a small distance to the implementation's own reference does not imply a small distance to the state that never stored the record. This paper installs a deletion interface into a pretrained language model and measures both di...

V. Ramesh · 3 citations
Jul 2026

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

It is shown, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut.

Saurabh Ranjan, K. Sokratous, Brian Odegaard · 0 citations
Preprint Aug 2026

Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution

The authors train 26M-parameter transformers on a synthetic language where recency and rarity are exactly coextensive, and separate them with a minimal causal edit that inverts one cue while holding the truth, token count and answer position fixed.

Yijun Liao, Fan-Wei Liang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.