A systematic shortcut audit of EmoPrefer using content-blind probes shows that the current scores can be reached without verifying either description against the video, and recommends source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations.
Abstract
Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track on EmoPrefer. Such benchmarks assume that predicting the preferred description requires grounded cross-modal understanding of the video. We conduct a systematic shortcut audit of EmoPrefer using content-blind probes. A simple logistic regression using only description length and generator identity, without processing the text, video, or audio, performs comparably to LoRA-finetuned 7B text and audio-visual judges (65.8 versus 66.8 WAF on EmoPrefer-V2). Generator identity is recoverable from description text with 99.5 percent accuracy, every candidate pair contrasts two distinct generators, and the human preference labels agree with a fold-exclusive per-generator win-rate prior on 66 percent of the evaluated pairs. When the human label conflicts with this prior, trained judges still follow the style prior on 63 to 80 percent of the pairs. On a length-matched subset that neutralizes verbosity bias, the tested media configurations yield no statistically significant improvement, while an ODIN-inspired diagnostic that decouples the style shortcut leaves its content head near chance. These results do not imply that human preferences are inherently stylistic or that the descriptions contain no emotional information. Instead, they show that the current scores can be reached without verifying either description against the video. We recommend source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations. Code is available at https://github.com/jiabingyang01/EmoPrefer-Audit.
We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annota...
Sofroniew et al. (2026) report that emotion concepts in Claude Sonnet 4.5 are represented as vectors whose geometry mirrors human affect psychology. We replicate the representational core of that study on the base pretrained model google/gemma-2-27b, inheriting every disclosed parameter, resolving unspecified steps by...
Experiments on the MER2026-EmoPrefer Challenge dataset and the error-augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence.
Zilong Huang, Junyi Peng, Junjie Li et al.· 0 citations
Emotion shapes credibility assessments, judicial decision-making, and perceptions of procedural justice, yet its reliable detection in courtroom settings remains a significant methodological challenge. This study provides a multimethod evaluation of emotion-recognition approaches in cross-linguistic legal discourse usi...
Jing-Yi Li, Di-Fu Shi, Huolingxiao Kuang et al.· Frontiers in Psychology· 0 citations
VIBE is introduced, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space and its core contribution is a measurement contract, which motivates entity-centered affective profiling as a documented practice.
A. Chetvergov, Alexander Evseev, Timofei Sivoraksha et al.· 0 citations
Emotion recognition in children’s drawings is difficult because affect is carried by sparse strokes, symbolic objects, and overall composition rather than by the stable appearance statistics of photographs. We built a reproducible four-class benchmark (Angry, Fear, Happy, Sad) on a single public corpus of 818 children’...
Hoonhee Lee, Min-woo Kim, Jaewon Kim et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.