This work introduces MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence.
Abstract
Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text representations. We introduce MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence. We also propose MemePIVOT, a local-global architecture for meme classification. MemePIVOT uses frozen CLIP features, unbalanced optimal transport to align words with image patches while allowing irrelevant evidence to remain unmatched, and an evidential fusion head to combine local grounding with global meme context under uncertainty. Experiments on HarMeme, PrideMM, and MemeCF show consistent gains over strong text-only, image-only, multimodal, and vision-language baselines. Cross-dataset and ablation results further show that explicit pivotal-evidence modeling improves robustness and contributes meaningfully beyond global multimodal representations. Our code and dataset are publicly available at https://github.com/AkshitSharma1/MemePIVOT
This paper presents M-NLE, a compact knowledge-augmented multitask model for jointly detecting offensive memes and generating natural language explanations, and suggests that knowledge-augmented explanation generation is a practical direction for more interpretable offensive meme detection.
Dibyanayan Bandyopadhyay, Baban Gain, Samrat Mukherjee et al.· IEEE Access· 0 citations
Harmful memes distribute risk signals across global scene context, localized visual cues, and embedded text, resisting unimodal detection. We propose TRIGUARD-LITE, a parameter-efficient framework that integrates three complementary views—global image semantics, five-region local crops, and embedded text features—throu...
Chuan Yu, Hui Deng· 2026 7th International Confe...· 0 citations
Detecting hate memes on social media presents a formidable challenge due to their multimodal and often subtle nature. The hateful intent typically emerges from a nuanced interplay between visual and textual elements, which existing methods often fail to capture by inadequately modeling these cross-modal correlations. T...
Lian-Song Zong, Qing-Chi Gui, Jie Wang et al.· IEEE Transactions on Computa...· 0 citations
Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the corre...
Zhen-Bin Wang, Lei Zhang, Li-Tuan Wang et al.· 0 citations
Misogynistic memes often convey harmful gender stereotypes through implicit interactions between images, text, and culturally specific references, making them difficult to detect with unimodal or general-purpose vision-language models. In this paper, we present a Chinese CLIP-based multimodal framework for misogynistic...
Mei-Lin Wu, Qin-Yue Li· Proceedings of the 28th Inte...· 1 citation
Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations, consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods.
Xiao-Ran Kou, Jing-Yi Wu, Peng Sun et al.· 0 citations