Skip to content
Conference

Do AI Agents Exhibit Greed in Shared Resource Environments?

Jul 2026 · International Conference on Artificial Intelligence Testing · pp. 27-30 · 0 citations · 17 references

Abstract

Large Language Models (LLMs) are increasingly used in simulations, either in academic settings for research studies or in industry for prototyping. Previous research has investigated the extent to which agents can mimic human behavior in socioeconomic settings; however, there is limited research on greedy decision-making by agents in simulated resource allocation environments. Furthermore, there is limited work on cross-model evaluation. Our research investigates the decision-making of ten agents across two different experimental conditions: one in which agents are able to communicate with other agents, and one in which emotional contexts are directly injected into the prompt. Based on a conceptual framework of greed, we find that agents predominantly exhibited greed-like behavior across all conditions. Interaction and self-reported social connection did not meaningfully influence the agents’ decision-making. We evaluated four different models: gpt-5-mini, gemini-3-flash-preview, claude-haiku-4-5-20251001, and grok-3-mini-fast-beta, and observed that while models differed in their self-reported connection scores, they did not differ significantly in greediness scores. The code and experimental artifacts are available at https://github.com/tiaL-ops/simCo.

View source

Similar papers

Preprint Aug 2026

AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions

This work presents AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration, a multi-agent forum for human--AI collaboration in scientific exploration that outperforms a centralized multi-agent debate baseline and shows that users value AgentPanel for perspective diversity and exploration support.

Zhiyao Cui, Qianyi Wang, Hao Yan et al. · 0 citations
Open access Jul 2026

Simulating strategic interactions with AI agents

This article introduces a framework for designing and running simulated experiments with LLM‐powered agents and applies the framework to the exploration–exploitation dilemma and shows that LLM‐based experiments reproduce patterns observed among human participants.

Matteo Tranchero, Cecil-Francis Brenninkmeijer, Arul Murugan et al. · 0 citations
#small language model Book Open access Sep 2026

Evaluating LLM Social Cognition Through Multi-Agentic Strategic Games

Large language models (LLMs) now power the reasoning core of intelligent virtual agents deployed across an expanding range of social settings, from tutoring students and supporting patients in healthcare, to mediating group discussions and representing humans in various social settings. Effective deployment demands social cognition, the capacity to model what others believe, detect deception, and coordinate strategic action under incomplete information. These capacities, exemplified in the social dynamics of the game Among Us, remain poorly characterized in current LLM evaluation frameworks. We introduce a strategic game arena that situates LLM agents in social deduction scenarios inspired by Among Us, requiring theory of mind, deception detection, and cooperative deliberation under uncertainty. We evaluate 19 open-weight models across 10,134 games and 289,614 utterances, testing both homogeneous and heterogeneous crews. Our experiments reveal three findings. First, crewmates voting through generative reasoning reach only \(50.4\% \pm 4.4\%\) F1 when identifying imposters, while a logistic regression classifier trained on the same discussion transcripts achieves \(85.3\%\) F1. Second, scaling model parameters yields a statistically significant but practically marginal improvement. Medium models (60–82B) reach \(52.8\%\) F1 against \(46.1\%\) for small models (7–20B), a 6.7 point gain (Mann–Whitney U, p = 1.5 × 10− 34). Third, agents fail to integrate evidence coherently during deliberation. Imposters self-incriminate in \(4.12\%\) of their statements, yet crewmates eject the confessing agent only \(33.8\%\) of the time. Crewmates reverse their stated suspect between consecutive rounds without new justification in \(41.9\%\) of cases. Sentiment remains uniformly neutral whether an agent is reporting a body or delivering a routine update. These gaps identify concrete limits in the social cognition of current LLM-powered agents and motivate architectural changes for virtual agents that must cooperate with humans. The source code and the live arena viewer are available at https://ufdatastudio.com/projects/agents-among-us.

Kevin Kurian, Kevin Scroggins, Emmanuel Dorley et al. · 0 citations
Book Open access Aug 2026

Performance Of Large Language Models As Hearthstone Agents

An LLM-driven Hearthstone agent is developed using the Sabberstone framework to evaluate several models, including GPT-4o, GPT-4o-mini, o3-mini, and GPT-5-mini, across multiple decks and prompting strategies, and results indicate that all evaluated LLMs outperform the random baseline, and GPT-5-mini achieves win rates close to the strongest numerical agents under the evaluation setting.

Christian Poglitsch, Philipp Bardakji, Johanna Pirker · 0 citations
#artificial intelligence Preprint Sep 2026

Interpreting and Steering LLM Agents for Social Simulations

Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents'internal representations into human-readable features, after which probes can reliably shift agents'behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.

Jia-Run Fan, Arul Murugan, Shreyas R. Krishnan et al. · 0 citations
Preprint Aug 2026

Social Behavior Among Autonomous AI: How Large Language Models Interact in Dynamic Networks

Cooperation is a cornerstone of human societies, enabling collective progress in dynamic and uncertain environments. With the advent of AI systems acting autonomously, it becomes crucial to understand not only human-AI cooperation but also AI-AI interactions in adaptive networks. In this work, we examine the interactions of AI using Large Language Models -- Mistral, Llama3, Gemma3, and Phi3 -- in a public goods game within dynamic network structures. Our experiments were conducted under single-model and mixed-model conditions across Watts-Strogatz (WS), Barabasi-Albert (BA), and Erdos-Renyi (ER) networks. We analyzed the impact of model architecture, network topology, and prompt design on cooperative behavior. Results show that Mistral and Llama3 offer high cooperation rates, while Phi3 shows defective tendencies. Additionally, the random structure of Erdos-Renyi networks dramatically improves cooperation. Prompt design also plays a key role; a society-benefits prompt leads to a higher cooperation level. These findings offer a preliminary framework for LLM-based simulations in adaptive social networks.

Narges Fardnia, Fatemeh Seyedin, Matthias Becker et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.