Skip to content

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Jul 2026 · arXiv.org · Vol abs/2607.28146 · 0 citations · 117 references
Computer Science

TL;DR

This work presents the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry, and introduces three novel metrics that isolate social deduction, reasoning, and deceptive consistency.

Abstract

As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.

View source

Similar papers

Jul 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

Marylou Fauchard, Florian Carichon, Margarida Carvalho et al. · 0 citations
Preprint Aug 2026

Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives

A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...

Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al. · 0 citations
#small language model Book Open access Sep 2026

Evaluating LLM Social Cognition Through Multi-Agentic Strategic Games

Large language models (LLMs) now power the reasoning core of intelligent virtual agents deployed across an expanding range of social settings, from tutoring students and supporting patients in healthcare, to mediating group discussions and representing humans in various social settings. Effective deployment demands soc...

Kevin Kurian, Kevin Scroggins, Emmanuel Dorley et al. · 0 citations
#natural language process... Open access Mar 2025

When Do Large Language Models Exhibit Unsolicited Deception?

Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are also better at prompted deception. But under what conditions do they deceive without instruction to do so? This study evaluates unsolicited deception produced by LLMs in a...

S. Taylor, B. Bergen · 11 citations
Preprint Aug 2026

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack. We designed exactly such a training signal -- a verifiable"survival"reward in which both the student's cited authorities an...

M. Kim · 0 citations
Preprint Aug 2026

Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs

This work formalises a necessary level-K distinguishability condition for strategic depth inference and builds a suite of novel game structures that meet this standard, and finds that model models maintain accurate strategic depth under recursive reasoning, with strong internal consistency between stated reasoning and...

Binchi Zhang, Atrisha Sarkar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.