Skip to content

LLM Unlearning Evaluation with TRIAGE

Danial Ataee Peter Triantafillou
Sep 2026 · 0 citations
Computer Science

TL;DR

TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.

Abstract

Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \emph{whether} a model appears to forget but provide limited insight into \emph{how} unlearning changes the model or affects related knowledge. We introduce \textit{TRIAGE} (\textit{Tripartite Representation-internal Introspection for Adjacency Gap Evaluation}), a benchmark-agnostic evaluation framework for characterizing these changes. TRIAGE uses diagonal approximations of the Fisher information and Hessian to measure changes in parameter sensitivity and local curvature, and utilizes a Forget / \emph{Adjacent-Retain} / \emph{Generic-Retain} partition to quantify an \emph{adjacency gap} in semantically related knowledge. Based on the magnitude and distribution of these changes, TRIAGE further classifies each algorithm's update as \emph{no-op}, \emph{partially localized}, \emph{collateral dominant}, or \emph{globally destructive}. Across 12 unlearning methods, four language models, and the WMDP, TOFU, and MUSE benchmarks, we find that methods with similar behavioral forgetting can produce substantially different internal changes and patterns of collateral damage. These signatures also vary across models and benchmarks, indicating that the effects of unlearning are not determined solely by the unlearning algorithm. TRIAGE can be applied alongside existing unlearning benchmarks to complement behavioral evaluation with a model-internal view of how unlearning reshapes the model's parameter space and affects retained knowledge.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal mov...

P. Mohanty, Hao-Ran Tang, Maggie Makar et al. · 0 citations
#machine learning Preprint Sep 2026

UnlearningSoup: Is Repeated Tuning Necessary for Large Language Model Unlearning?

Whether unlearning runs exhibit exploitable structure in weight space is investigated, and it is observed that models from different runs still lie in a shared evaluation-performance basin, suggesting that stronger models may be recovered through an unlearning-tailored soup strategy, reducing the need for repeated tuni...

Pu-Ning Yang, Qi-Zhou Wang, Jun-Chi Yu et al. · 0 citations
Preprint Aug 2026

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

It is argued that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning.

Sahil Kale, Ian G. Harris · 0 citations
Preprint Aug 2026

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

ADU is presented, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling, and achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks.

Xun-Lei Chen, Qi-Rui Ye, Yuang Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UNBIND: UNlearning By INference-time Directional Steering for Code LLMs

Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patt...

Zhengyang Shan, Jia-Yu Xin, Yan-Jun Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

This work proposes Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score and provides an empirical path toward quantization-resilient unlearning.

Ravi Ranjan, O. Kotevska, Agoritsa Polyzou · 1 citation

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.