Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patt...
Zhengyang Shan, Jia-Yu Xin, Yan-Jun Lin et al.· 0 citations
OBLIVION, a controlled benchmark and defense harness for revoked-skill resurrection, and results support workflow-level evaluation beyond checking explicit skill entries support workflow-level evaluation beyond checking explicit skill entries.
Zhengyang Shan, Xuancheng Qian, Jiayu Xin et al.· 0 citations
Mask2Shield (M2S), a masked-forward alignment method that trains a model under this functional pruning procedure, reduces successful recomputed pruning attacks from 80--279 to 1--44 out of 313 prompts while generally preserving four capability benchmarks.