Skip to content
Preprint

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

Aug 2026 · 0 citations · 38 references
Computer Science

TL;DR

The AdaPop (Adaptive Popularity) method is proposed, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy, and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch.

Abstract

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can de...

Ravi Ranjan, O. Kotevska, Agoritsa Polyzou · 0 citations

On-the-go Forgetting without Explicit Unlearning via ERASE

Existing unlearning approaches typically rely on post hoc weight adaptation or distillation, leading to duplicated memory costs, degraded generalization, and limited scalability. In this work, we introduce ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses...

Kushal Chakrabarti, Mayank Baranwal · 0 citations
#machine learning Preprint Sep 2026

Test-Time Unlearning via Sparse Autoencoder

Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify model weights via gradient ascent and its advances. While effective on certain benchmarks, these weight-based approaches exhibit a sharp forget-utility trade-off, where...

Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi et al. · 0 citations
2026

DeepU: Deeper Granular Within-Layer Machine Unlearning

Machine unlearning (MU) aims to remove the influence of selected data from trained models, offering an efficient alternative to full retraining. With the rise of increasingly stringent privacy regulations, including the right to be forgotten, machine learning models must incorporate mechanisms that ensure compliance wh...

Anudeep Vurity, Zhi-Sheng Yan, Massimiliano Albanese · 0 citations
#artificial intelligence Preprint Aug 2026

GRACE:Gradient-guided Coreset Selection for LLM Unlearning

GRACE, a gradient-guided coreset selection method that constructs both forget and retain sets for LLM unlearning, improves model utility while maintaining comparable forget quality, with particularly consistent gains over prior gradient-based selection methods.

Praveen Bushipaka, Andrea D'Angelo, Lucia C. Passaro et al. · 0 citations
Preprint Aug 2026

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

ADU is presented, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling, and achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks.

Xun-Lei Chen, Qirui Ye, Yuang Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.