Skip to content

Selective Knowledge Edit Reversal via Gated Singular Vector Shrinkage

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.

Abstract

Knowledge editing provides an efficient way to update factual knowledge in large language models. However, malicious edits may introduce safety risks, making it necessary to reverse undesirable editing effects. Existing reversal methods for parameter-modifying edits mainly focus on global removal, which may also erase beneficial edits that should be preserved. In this paper, we study selective reversal of edited knowledge, where the goal is to reverse targeted edited facts while preserving the remaining edited facts. Based on the hypothesis that each edit is sparsely encoded within the dominant subspace of the edited matrix, we propose a spectral-based reversal framework that locates edit-sensitive components within the dominant singular subspace of edited weights. Experiments across multiple settings demonstrate the effectiveness of our method in reversing selected edits while preserving unrelated edited facts. These results suggest that different edits are sparsely encoded within dominant singular components and can be separable when the number of edits is moderate, making selective spectral reversal a promising direction for locating edit-specific components and repairing edited language models.

View source

Similar papers

#machine learning Preprint Sep 2026

Constrained Edit Fields for Training-Free Flow Editing

Constrained Edit Fields (CEF) achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment.

Jing-Xuan Kang, Yin-Song Wang, Che Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ALOE: Semantically Addressed Low-Rank Operators for Knowledge Editing

Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates to...

Zeyan Li, Hu Xu, Jian-Feng Xu · 0 citations
#natural language process... Preprint Sep 2026

ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective

Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forgetting and degradati...

Rui Liu, Chen-Heng Zhang, Hao-Xuan Li et al. · 0 citations
Book Open access Aug 2026

Directional Time Series Editing via Retrieval-Guided Jacobian-Vector Inference

A new TSE setting for continuous, magnitude-aware condition transitions is introduced and JAVELIN, a retrieval-guided framework for directional editing via JAcobian-VEctor Latent INference is proposed, enabling precise, content-preserving edits without retraining the generative model.

Yifan Bao, Yihao Ang, Qiang Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation

KLOD is proposed, a bounded and distribution-preserving objective for fine-tuning-based knowledge editing that separates the intended target update from distributions that should remain stable and substantially mitigates locality degradation while maintaining high edit reliability.

Hojun Jeong, Gyunyeop Kim, Sangwoo Kang · 0 citations
#machine learning Preprint Sep 2026

Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery

Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities. Understanding the internal computations that support these adaptations is therefore becoming increasingly important for continual model evolution. Sparse autoencoders...

Hendrik Droste, C. M. Adriano, Kathrin Korte et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.