Skip to content
Book Open access

CtrlYoMel@CC-MMD 2026: A Chinese CLIP-Based Multimodal Framework for Misogynistic Meme Detection

Oct 2026 · Proceedings of the 28th International Conference On Multimodal Interaction · 1 citation · 10 references

Abstract

Misogynistic memes often convey harmful gender stereotypes through implicit interactions between images, text, and culturally specific references, making them difficult to detect with unimodal or general-purpose vision-language models. In this paper, we present a Chinese CLIP-based multimodal framework for misogynistic meme detection on the Chinese partition of the Cross-Cultural Misogynistic Meme Detection benchmark. We encode meme images and Chinese text into a shared semantic space, apply L2 normalization, and concatenate the representations for classification with a lightweight multilayer perceptron. We jointly fine-tune the Chinese CLIP backbone and classifier and use weighted cross-entropy to address class imbalance. On the official test set, our system achieves an accuracy of 0.95294 and a macro-F1 score of 0.94429, ranking first among six valid teams. Baseline and ablation results confirm the benefits of multimodal fine-tuning, normalization, concatenation-based fusion, and dropout. Error analysis highlights remaining challenges involving visual metaphor, culturally situated meanings, and confusion between general aggression and gender-targeted hostility, motivating more culturally grounded and target-aware multimodal reasoning.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.