Skip to content
Preprint

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

The Forget-Retain Alignment Gap is introduced, a training-free predictor that scores an update's forget-retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance, suggesting that weight selectivity better explains robustness than distance alone.

Abstract

Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget-critical weights while sparing retain-critical ones. We introduce the Forget-Retain Alignment Gap (FRAG), a training-free predictor that scores an update's forget-retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance. Building on the forget-critical, retain-sparing principle, Forget-Retain Pruning (FRP) improves relearning robustness. Our results suggest that weight selectivity better explains robustness than distance alone. Code is available at https://github.com/Yi1-Chen/FRAG.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

This work proposes Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score and provides an empirical path toward quantization-resilient unlearning.

Ravi Ranjan, O. Kotevska, Agoritsa Polyzou · 1 citation
#machine learning Preprint Sep 2026

When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data

Machine unlearning aims to remove the influence of designated training data while preserving model utility, but its behavior on tabular data remains underexplored. This gap is important because tabular prediction is widely used in high-stakes domains and is increasingly adapted to language models through record seriali...

Zijie Liu, Jinhao Duan, Bing-Qi Shang et al. · 0 citations
Preprint Aug 2026

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

The AdaPop (Adaptive Popularity) method is proposed, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy, and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch.

Anna Borisiuk, A. Savchenko, Alexander Panchenko et al. · 0 citations
Preprint Aug 2026

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

ADU is presented, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling, and achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks.

Xun-Lei Chen, Qi-Rui Ye, Yuang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

CONfession-to-Forget-Set (CONFS), a data-blind framework that constructs model-aligned forget sets by eliciting and formalizing the model's memorized knowledge, approaches Gold-standard performance on several metrics and achieves a competitive forgetting-utility balance, while preserving utility better than other data-...

Miso Kim, Georu Lee, Seungwon Jeong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.