Skip to content

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

Jun 2026 · arXiv.org · Vol abs/2606.29689 · 1 citation · 50 references
Computer Science

TL;DR

It is argued that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique, and that reference-based similarity gives a misleading picture.

Abstract

Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather than the written critiques people actually give. We ask whether MLLM critiques are close to human ones, scoring eight open-weight MLLMs from $7$B to $397$B, plus GPT-5.5, against multiple ranked human critiques for each of $1{,}227$ \texttt{r/photocritique} posts under eight prompt conditions. Reference-based similarity gives a misleading picture. In absolute terms the stricter lexical and learned metrics align only weakly with human critiques while a coarse embedding cosine reports broad topical overlap, yet requesting shorter critiques raises those scores and withholding the image barely changes them: the similarity reflects length, the post text, and a stable critiquing style more than image-specific observation. An LLM judge sharpens the question rather than settling it: in the primary condition all four judges prefer the frontier models'critiques to the human ones, but on the $7$--$8$B models they diverge wildly, from $9\%$ to $81\%$ preference on identical pairs. Asked instead how similar each pair is in substance, those judges and two human annotators agree, rating every model between $1.81$ and $2.59$ on a $1$--$5$ scale, close to ``mostly different''. Behaviorally, the models diverge in ways the scores do not surface: they cover nearly every aesthetic aspect where humans are selective and repeat themselves across critiques of one photo, even when prompted to write at human length. We argue that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique.

View source

Similar papers

#machine learning Preprint Sep 2026

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short s...

Rui-Chen Zheng, Yi-He Wang, Fabrice Y. Harel-Canada et al. · 0 citations
Open access Sep 2026

Reasoning Together: Designing and Evaluating MLLM Team Strategies for Multimodal Quiz Questions

Multimodal large language models (MLLMs) often struggle with open-ended questions requiring the integration of visual evidence, indirect textual clues, and background knowledge. We investigate whether team-based inference improves performance on Russian-language multimodal What? Where? When? questions. We introduce a n...

A. Kotelnikova, V. Byzov, M. Dolzhenkova et al. · 0 citations
#small language model Open access Sep 2026

Semantic anchoring with concise ideal answers outperforms unstructured full course materials as context for multi-LLM automated grading of open-ended questions

Large language models (LLMs) are increasingly used to grade open-ended student responses, yet the role of contextual input in this process remains poorly understood. This study compares three context conditions for multi-LLM automated grading: no context, full course materials, and instructor-defined ideal answers...

Jorge Cisneros-González, Natalia Gordo-Herrera, Iván Barcia-Santos et al. · 0 citations
#artificial intelligence Preprint Sep 2026

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

A benchmark for evaluating whether LLMs can recover situated pragmatic meanings in Chinese online comments is introduced, and case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.

Yin-Hong Shi, Jun-Jie Ma, Emma Jiren Wang et al. · 0 citations
Preprint Aug 2026

Can Open-Weight Models Compete on Financial Text Comprehension?

Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495...

J. Spörer · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.