Book
Open access
Aug 2026
SAFT: Safety-Preserving Adaptation via Fine-Tuning Transfer for Large Language Models
SAFT (Safety-preserving Adaptation via Fine-tuning Transfer), a safety-preserving adaptation framework that decouples task learning from alignment preservation by learning a safety-guided task update on the paired pretrained base model, rectifying task gradients to avoid conflicting directions with respect to a safety objective, and then transferring the update to the frozen instruction model via parameter-space grafting.
Zhiwen Ruan, Yan Yang, Zhuocheng Liang et al.
· Proceedings of the 32nd ACM... · 0 citations