Preprint
Aug 2026
MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
Experiments show that MMAligner raises the average refusal rate on unsafe multimodal inputs to 99% with less than 2% utility degradation and minimal training data, substantially improving the safety-utility trade-off over existing baselines.
Shenyi Zhang, Keyan Guo, Zihao Wang et al.
· 0 citations