Abstract Concept Grounding through Safety Alignment in Multimodal Large Language Models
Abstract
Abstract concepts such as harmfulness, legality, privacy, and fraud present a major challenge for multimodal large language models (MLLMs), as they require reasoning across visual, linguistic, and contextual information rather than direct perception. This paper investigates three questions: whether supervised alignment improves abstract concept grounding, how preference-based optimisation influences concept discrimination, and whether a two-stage SFT→DPO strategy outperforms individual alignment methods. We evaluate Qwen2.5-VL-7B-Instruct on MM-SafetyBench, which covers 13 categories of safety-related abstract concepts. Using benchmark-derived SFT and DPO datasets, we compare supervised fine-tuning (SFT), Direct Preference Optimisation (DPO), and a sequential SFT→DPO pipeline. Results provide preliminary behavioural evidence consistent with improved discrimination of safety-related abstract concepts. SFT achieves the strongest reduction of unsafe outputs, while DPO demonstrates competitive performance while maintaining behaviour learned through preference optimisation. The combined SFT→DPO framework provides the best overall balance between safety and response quality. Our findings suggest that multimodal safety alignment may provide a useful lens through which abstract concept processing can be investigated, where concepts such as harm, deception, and legality emerge through cross-modal reasoning. Our framework offers a reproducible approach for studying abstract concept grounding in multimodal systems.