Skip to content

Small Cues, Big Consequences: Learning Pivotal Cues for Multimodal Meme Classification

Sep 2026 · 1 citation · 55 references
Computer Science

TL;DR

This work introduces MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence.

Abstract

Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text representations. We introduce MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence. We also propose MemePIVOT, a local-global architecture for meme classification. MemePIVOT uses frozen CLIP features, unbalanced optimal transport to align words with image patches while allowing irrelevant evidence to remain unmatched, and an evidential fusion head to combine local grounding with global meme context under uncertainty. Experiments on HarMeme, PrideMM, and MemeCF show consistent gains over strong text-only, image-only, multimodal, and vision-language baselines. Cross-dataset and ablation results further show that explicit pivotal-evidence modeling improves robustness and contributes meaningfully beyond global multimodal representations. Our code and dataset are publicly available at https://github.com/AkshitSharma1/MemePIVOT

View source

Similar papers

Review Open access 2026

M-NLE: Knowledge-Augmented Multitask Learning for Offensive Meme Detection and Explanation Generation

This paper presents M-NLE, a compact knowledge-augmented multitask model for jointly detecting offensive memes and generating natural language explanations, and suggests that knowledge-augmented explanation generation is a practical direction for more interpretable offensive meme detection.

Dibyanayan Bandyopadhyay, Baban Gain, Samrat Mukherjee et al. · 0 citations
Conference Aug 2026

TriGuard-Lite: Parameter-Efficient Harmful Meme Recognition with Harm-aware Local Regions and Gated Multimodal Fusion

Harmful memes distribute risk signals across global scene context, localized visual cues, and embedded text, resisting unimodal detection. We propose TRIGUARD-LITE, a parameter-efficient framework that integrates three complementary views—global image semantics, five-region local crops, and embedded text features—throu...

Chuan Yu, Hui Deng · 0 citations
Oct 2026

Context-Aligned Latent Enhancement for Multimodal Hate Meme Detection

Detecting hate memes on social media presents a formidable challenge due to their multimodal and often subtle nature. The hateful intent typically emerges from a nuanced interplay between visual and textual elements, which existing methods often fail to capture by inadequately modeling these cross-modal correlations. T...

Lian-Song Zong, Qing-Chi Gui, Jie Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the corre...

Zhen-Bin Wang, Lei Zhang, Li-Tuan Wang et al. · 0 citations
Book Open access Oct 2026

CtrlYoMel@CC-MMD 2026: A Chinese CLIP-Based Multimodal Framework for Misogynistic Meme Detection

Misogynistic memes often convey harmful gender stereotypes through implicit interactions between images, text, and culturally specific references, making them difficult to detect with unimodal or general-purpose vision-language models. In this paper, we present a Chinese CLIP-based multimodal framework for misogynistic...

Mei-Lin Wu, Qin-Yue Li · 1 citation
Preprint Sep 2026

Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis

Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations, consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods.

Xiao-Ran Kou, Jing-Yi Wu, Peng Sun et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.