CtrlYoMel@CC-MMD 2026: A Chinese CLIP-Based Multimodal Framework for Misogynistic Meme Detection
Abstract
Misogynistic memes often convey harmful gender stereotypes through implicit interactions between images, text, and culturally specific references, making them difficult to detect with unimodal or general-purpose vision-language models. In this paper, we present a Chinese CLIP-based multimodal framework for misogynistic meme detection on the Chinese partition of the Cross-Cultural Misogynistic Meme Detection benchmark. We encode meme images and Chinese text into a shared semantic space, apply L2 normalization, and concatenate the representations for classification with a lightweight multilayer perceptron. We jointly fine-tune the Chinese CLIP backbone and classifier and use weighted cross-entropy to address class imbalance. On the official test set, our system achieves an accuracy of 0.95294 and a macro-F1 score of 0.94429, ranking first among six valid teams. Baseline and ablation results confirm the benefits of multimodal fine-tuning, normalization, concatenation-based fusion, and dropout. Error analysis highlights remaining challenges involving visual metaphor, culturally situated meanings, and confusion between general aggression and gender-targeted hostility, motivating more culturally grounded and target-aware multimodal reasoning.