Skip to content

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

Jun 2026 · arXiv.org · Vol abs/2606.23092 · 0 citations · 39 references
Computer Science

TL;DR

PIVOTS is introduced, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research and examines how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions.

Abstract

Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs'ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models'ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions. Project page: https://ciossayin.github.io/pivots-bench/

View source

Similar papers

Preprint Aug 2026

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

The results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships.

Zhanna Mukhametsharip, Vera Demberg, Varsha Suresh Saarland University et al. · 0 citations
#artificial intelligence Preprint Sep 2026

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

A benchmark for evaluating whether LLMs can recover situated pragmatic meanings in Chinese online comments is introduced, and case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.

Yin-Hong Shi, Jun-Jie Ma, Emma Jiren Wang et al. · 0 citations
Preprint Aug 2026

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...

Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al. · 1 citation
#computer vision Preprint Aug 2026

SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos

SocialReasonBench is introduced, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives and develops a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guide...

Zhe-Yu Huang, Zijing Shi, Hao-Zhe Luo et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.