These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.
Abstract
Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.
The results show that group identity shapes how LLMs aggregate information across agents, independently of its correctness, and identify a manipulation surface for multi-agent AI systems.
A dual-level evaluation framework to assess LLM-based agents at both the individual and collective levels is proposed, finding that while agents capture broad partisan orientations, they underestimate within-group variability and reproduce stereotypical ideological biases.
M. Al Ali, Filip Mihai Muntean, Lucia Donatelli et al.· International Conference on...· 1 citation
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer...
The results showed that social exclusion reduced conformity to AI-generated opinions, but not to human opinions, while task type did not significantly moderate the effect of opinion source, suggesting that conformity to AI may be shaped by mechanisms that differ from those underlying traditional interpersonal conformit...
Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physic...
This work proposes A-B-D to infer traits bottom-up from behavioral data (B-data), namely how agents act on their environment and communicate with users, as recorded in existing trajectories, and offers a new lens for understanding AI personality.
Hao-Kai Zhao, Jie Gao, Yunze Xiao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.