TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization, is proposed.
Mansi, Avinash Kori, Francesca Toni et al.· arXiv.org· 2 citations
Eval-unlearn is an open-source Python library providing a unified, reproducible benchmarking framework for concept unlearning in T2I Diffusion models, integrating twelve published unlearning techniques spanning fine-tuning, closed-form model editing, and inference-time intervention.
Mansi, Nikhil Raghavan, Zi-Xia Huang et al.· 0 citations
Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence over a finite set of queries and leave residual leakage over the broader prompt space lar...
Mansi, Luca Marzari, Francesco Leofante· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.