MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints, is introduced.
Abstract
Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
Different capacities for mentalization across LLMs are demonstrated, and cognitive computational modeling is highlighted as a formal method for assessing comparative intelligence across humans and machines.
Aamir Sohail, Xintong Zhong, Arkady Konovalov et al.· 0 citations
A central challenge in understanding joint action is explaining how individuals achieve successful coordination in dynamic, real-world settings. Although coordination is thought to depend on anticipating future events, including the actions of others, the precise contribution of such predictive processing remains uncle...
P. Putra, Fumihiro Kano· Communications Psychology· 0 citations
Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produces response dynamics that action-only formulations leave unmodeled, such as verbal production, intera...
Jiawen Kang, Dongrui Han, Xi-Xin Wu et al.· 0 citations
MindEvolve is introduced, an autonomous workflow designed to predict behavior in social interactions by generating interpretable symbolic models of cognition, and provides a roadmap for advancing LLM-based cognitive modeling toward human-expert-level theory construction.
Ying-Ying Ye, You-Le Fang, Xiao-Xue Gao et al.· bioRxiv· 0 citations
Theory of mind is central to human social behavior, yet empirical studies rely largely on biased measures with limited ecological validity. We propose and evaluate the use of mental state language as a spontaneous indicator of theory of mind expressed in social interaction. Among research participants (334 undergraduat...
Chloe C. Hudson, Louis Hickman, Si-Yi Liu et al.· Assessment (Odessa, Fla.)· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.