Skip to content

Popular Knowledge Propagates More Errors in LLM Knowledge Updating

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

The results reveal a pattern distinct from prior findings on long-tail vulnerability during acquisition and retention: among facts that models already answer correctly, those associated with highly connected entities are more likely to be corrupted by neighboring updates, and updates to such facts propagate errors more broadly.

Abstract

Updating a language model's knowledge through fine-tuning is essential for keeping its outputs current, yet can also induce factual forgetting and new hallucinations. Prior work shows that long-tail knowledge is harder to acquire and newly memorized long-tail facts are difficult to retain during later fine-tuning. We study a complementary question: among facts that a model has encoded correctly, which are most vulnerable to collateral corruption during other updates? To investigate this question under a realistic factual distribution, we construct a large-scale graph FACTPROP of verified Wikipedia facts by linking triples that share head or tail entities, thereby preserving connections among factual knowledge. We fine-tune models on factual statements and measure correct-to-incorrect facts after each update. Our results reveal a pattern distinct from prior findings on long-tail vulnerability during acquisition and retention: among facts that models already answer correctly, those associated with highly connected entities are more likely to be corrupted by neighboring updates, and updates to such facts propagate errors more broadly. Structural popularity therefore predicts both vulnerability and downstream damage. Inspired by this finding, we propose Popularity-based Anchoring (PopAnchor), a lightweight rehearsal strategy that preserves a small set of popular facts and reduces forgetting.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates to...

Zeyan Li, Hu Xu, Jian-Feng Xu · 0 citations
#artificial intelligence Preprint Sep 2026

Probing Factual Knowledge Transfer with Training Data Interventions

The results show that fact transfer is very limited: under the strictest removal condition, a large majority of English-acquired facts fail to transfer into Persian, and it is shown that sentence-level co-occurrence removal is insufficient to eliminate fact signal.

Romina Oji, Marc Braun, Marcel Bollmann et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Memory vs. Context? Influential Factors of Factual Recall in Language Models

We reproduce and stress-test the work of Yu et al. (2023), who characterize how language models (LMs) arbitrate between memorized knowledge and contradictory in-context statements. We replicate their world-capitals experiments on 31 models spanning Pythia, GPT-2, Qwen3, and Ministral families, including base and post-t...

Guilhem Fouilhé, Nicholas Asher, Philippe Muller · 0 citations
Preprint Aug 2026

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

HPSE is proposed, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere.

Tianci Liu, Zi-Han Dong, Tian-Chun Li et al. · 0 citations
#machine learning Preprint Sep 2026

Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths

While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact. To achieve true fo...

Jia-Lu Wang, Pei-Zhi Niu, Hao-Teng Yin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.