Jul 2026· Journal of Cybersecurity and Privacy· Vol 6, pp. 121· 0 citations· 35 references
TL;DR
This work proposes Text-to-Unlearn, a novel framework that selectively unlearns concepts from pre-trained GANs using only text prompts, enabling feature and identity unlearning, as well as fine-grained tasks such as expression and multi-attribute removal in models trained on human faces.
Abstract
State-of-the-art generative models exhibit powerful image-generation capabilities, raising ethical and legal challenges for service providers. Consequently, Content Removal Techniques (CRTs) have emerged to control outputs without requiring full retraining. However, the problem of unlearning in Generative Adversarial Networks (GANs) remains largely unexplored. We propose Text-to-Unlearn, a novel framework that selectively unlearns concepts from pre-trained GANs using only text prompts, enabling feature and identity unlearning, as well as fine-grained tasks such as expression and multi-attribute removal in models trained on human faces. Our approach leverages natural language descriptions to guide unlearning without additional datasets or supervised finetuning, offering a scalable solution. To evaluate the effectiveness of our method, we introduce an automated unlearning assessment method using state-of-the-art image–text alignment metrics and propose a new metric: degree of unlearning. Additionally, we assess robustness by introducing adversarial attacks to subvert unlearning. Our results demonstrate that Text-to-Unlearn achieves robust unlearning, resisting adversarial attempts to recover erased concepts while preserving model utility. To our knowledge, this is the first cross-modal unlearning framework for GANs, advancing the management of generative model behavior.
The results reveal that attention-space regulation offers a considerably more promising path to safer diffusion transformer based image generation than the existing concept erasing mechanism.
The paradigms of the generative adversarial networks (GANs), variational autoencoders (VAEs), transformer-based designs, and diffusion models, with the last one representing the state of the art in image generation models are discussed.
DiSCO is proposed, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals, and can be readily applied to any text-to-image system without necessitating any changes to the model itself.
Tong Zhang, M. Alfarra, Carlos Hinojosa et al.· 0 citations
A lightweight Text Encoder Alignment framework that fine-tunes only the text encoder while keeping the generative backbone fully frozen, and achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts.
Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal comprehension capabilities, achieving state-of-the-art performance across various vision-language tasks. However, their performance drops significantly when facing adversarial attacks on the visual encoder. To alleviate this issue, existing ap...
Bo-Yu Wang, Zi-Wen He, Xin-Jue Hu et al.· ACM Transactions on Multimed...· 0 citations
This work introduces PRMU, a benchmark for evaluating corpus-free multimodal unlearning under realistic person-centric deletion requests, and introduces Similarity-Gated Projection Editing (SGPE), a lightweight corpus-free unlearning baseline with knowledge displacement, protected parameter-space editing, and locality-...
Hua-Feng Chen, Yueming Lyu, Ziyuan Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.