Jun 2026· arXiv.org· Vol abs/2606.29689· 1 citation· 50 references
Computer Science
TL;DR
It is argued that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique, and that reference-based similarity gives a misleading picture.
Abstract
Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather than the written critiques people actually give. We ask whether MLLM critiques are close to human ones, scoring eight open-weight MLLMs from $7$B to $397$B, plus GPT-5.5, against multiple ranked human critiques for each of $1{,}227$ \texttt{r/photocritique} posts under eight prompt conditions. Reference-based similarity gives a misleading picture. In absolute terms the stricter lexical and learned metrics align only weakly with human critiques while a coarse embedding cosine reports broad topical overlap, yet requesting shorter critiques raises those scores and withholding the image barely changes them: the similarity reflects length, the post text, and a stable critiquing style more than image-specific observation. An LLM judge sharpens the question rather than settling it: in the primary condition all four judges prefer the frontier models'critiques to the human ones, but on the $7$--$8$B models they diverge wildly, from $9\%$ to $81\%$ preference on identical pairs. Asked instead how similar each pair is in substance, those judges and two human annotators agree, rating every model between $1.81$ and $2.59$ on a $1$--$5$ scale, close to ``mostly different''. Behaviorally, the models diverge in ways the scores do not surface: they cover nearly every aesthetic aspect where humans are selective and repeat themselves across critiques of one photo, even when prompted to write at human length. We argue that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique.
LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short s...
Rui-Chen Zheng, Yi-He Wang, Fabrice Y. Harel-Canada et al.· 0 citations
Multimodal large language models (MLLMs) often struggle with open-ended questions requiring the integration of visual evidence, indirect textual clues, and background knowledge. We investigate whether team-based inference improves performance on Russian-language multimodal What? Where? When? questions. We introduce a n...
A. Kotelnikova, V. Byzov, M. Dolzhenkova et al.· Computers· 0 citations
Large language models (LLMs) are increasingly used to grade open-ended student responses, yet the role of contextual input in this process remains poorly understood. This study compares three context conditions for multi-LLM automated grading: no context, full course materials, and instructor-defined ideal answers...
A benchmark for evaluating whether LLMs can recover situated pragmatic meanings in Chinese online comments is introduced, and case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.
Yin-Hong Shi, Jun-Jie Ma, Emma Jiren Wang et al.· 0 citations
Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495...
These findings show that LLM-assisted peer review changes the functional composition of review text, making it important to distinguish LLM-amplified critique from areas requiring human prioritization and accountable judgement.
YunHong Yang, Mike Thelwall, Guo-Xiu He· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.