Skip to content

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Jul 2026 · arXiv.org · Vol abs/2607.27798 · 1 citation · 40 references
Computer Science

TL;DR

KAR, an entity-guided retrieval baseline built on CultureBase is introduced and MemeBench is introduced, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures, to reveal whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

Abstract

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

View source

Similar papers

#computer vision Preprint Sep 2026

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

This work introduces MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes.

Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir et al. · 0 citations
Preprint Aug 2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

PoVisLE is introduced, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context.

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn et al. · 1 citation
#small language model Preprint Aug 2026

Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

CMPM, a Chinese Multi-Panel Meme benchmark with 1,214 annotated samples covering five structural types, ordering dependency, panel-order constraints, and optional comment context, is introduced and results indicate that canonical-display accuracy is not by itself evidence of order understanding.

Haihan Li, Hai-Hao Li, Zheng-Jie Xu et al. · 0 citations
#natural language process... Preprint Sep 2026

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

A language model normally begins training with random word embeddings: whatever'banana'means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens rece...

Lisa Bylinina · 0 citations
Preprint Aug 2026

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

The results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships.

Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh Saarland University et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.