Jun 2026· arXiv.org· Vol abs/2606.23092· 0 citations· 39 references
Computer Science
TL;DR
PIVOTS is introduced, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research and examines how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions.
Abstract
Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models'ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions. Project page: https://ciossayin.github.io/pivots-bench/
The results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships.
Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh Saarland University et al.· 0 citations
This work proposes CaRGo-T (Causal Reasoning Graph-of-Thought), a reasoning framework that represents the causal and contextual relationships underlying multimodal humor as a lightweight graph-based reasoning structure.
Abhilash Nandy, Rahul Seetharaman, Aman Bansal et al.· 0 citations
A benchmark for evaluating whether LLMs can recover situated pragmatic meanings in Chinese online comments is introduced, and case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.
Yin-Hong Shi, Jun-Jie Ma, Emma Jiren Wang et al.· 0 citations
CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context.
Bo Zeng, Lin-Feng Gao, Pei-Qing Lin et al.· 0 citations
EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...
Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al.· 1 citation
SocialReasonBench is introduced, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives and develops a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guide...
Zhe-Yu Huang, Zijing Shi, Hao-Zhe Luo et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.