Skip to content

LLMs Trust Their Own: Identity-Dependent Conformity in Multi-Agent Systems

Sep 2026 · 1 citation · 58 references
Computer Science

TL;DR

The results show that group identity shapes how LLMs aggregate information across agents, independently of its correctness, and identify a manipulation surface for multi-agent AI systems.

Abstract

Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs'responses depend on the social identity of other agents, beyond the effect of their consensus. We construct judgment tasks with a single correct answer, and place models in a multi-agent setting where they receive incorrect answers from other agents whose social identities (AI or human, model family, or an arbitrary minimal group) are either shared with or distinct from their own. Across 12 open-weights models and nine tasks, we find a bidirectional effect of group identity on conformity to incorrect answers: in-group consensus increases conformity (in-group favoritism), whereas out-group consensus decreases it (out-group divergence). Unlike humans, for whom one ally breaking the consensus sharply reduces conformity, models are unmoved by an ally from the majority's group. Worse, a correct ally from the opposing group intensifies this bidirectional effect. Chain-of-Thought reasoning suppresses most of these effects, yet an in-group ally still reduces conformity to an incorrect out-group majority. Labeling peers as safety-aligned shifts overall conformity but leaves in-group favoritism and out-group divergence intact. These results show that group identity shapes how LLMs aggregate information across agents, independently of its correctness, and identify a manipulation surface for multi-agent AI systems.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems

Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agents prefer interacting with others carryin...

Xavier Del Giudice, Alessio Palma, Matteo Migliarini et al. · 0 citations
#natural language process... Preprint Sep 2026

AI agents reshape consensus formation in human groups

As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated r...

Lin Chen, Ziyi Liu, Xia Hu et al. · 0 citations
#machine learning Preprint Sep 2026

Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation

Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both...

Cong-Ling Li, Cheng Chen, Thomas Fung et al. · 0 citations
#small language model Preprint Aug 2026

Predicting the scale limits of social mechanisms in agent societies

An audit is introduced that predicts a mechanism's fate as a population grows, asking how often the mechanism can act, whether agents use the information it supplies, and whether the measurement itself creates apparent scale effects.

Zengqing Wu, Chuan Xiao · 0 citations
#artificial intelligence Preprint Sep 2026

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory t...

Carolina Fortuna, Blaž Bertalanič · 0 citations
Preprint Sep 2026

Collective Regimes in Multi-Agent LLMs under Reasoning Effort and Communication Topology

Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such measures do not reveal how agreement is org...

Machiko Hirota, Akshara Nadayanur Sathis Kanna, Ujwal Kumar et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.