Skip to content
Review

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Jul 2026 · arXiv.org · Vol abs/2607.19011 · 0 citations · 121 references
Computer Science

TL;DR

This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier, and synthesizes benchmark design, evaluation protocols, and modeling paradigms based on multimodal alignment, evidence-grounded reasoning, and controlled generation.

Abstract

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

Though absolute performance remains low, finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance.

Akhila Yerukola, Fabrice Y. Harel-Canada, Simran Khanuja et al. · 0 citations
Jul 2026

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

KAR, an entity-guided retrieval baseline built on CultureBase is introduced and MemeBench is introduced, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures, to reveal whether a...

Weihang Wang, Kainan Tu, Jielei Zhang et al. · 1 citation
Preprint Aug 2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

PoVisLE is introduced, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context.

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.