Skip to content
Open access

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment.

Sep 2026 · Proceedings of the National Academy of Sciences of the United States of America · Vol 123 37, pp. e2600126123 · 0 citations · 25 references
Medicine

Abstract

Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments that others can understand and publicly accept. Using the Deliberative Reason Index (DRI), this study evaluates 60 LLMs against human deliberation across nine policy scenarios. Only four models consistently exceed the permutation-based null benchmark for alignment with human patterns of reason-giving. Most models fall short: their reason-preference structures rarely clear this threshold, even though their outputs can still appear coherent and persuasive. Yet outputs can appear reasonable even when this alignment is absent. The observed gap between surface plausibility and deliberative coherence urges caution: deploying LLMs in governance contexts requires prior assessment of their deliberative reasoning capacity, not just their surface outputs.

Read PDF

Similar papers

Preprint Aug 2026

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer e...

Maurice Flechtner · 2 citations

JudgmentLens: Human-AI Sensemaking of Complex Legal Judgments

Judicial judgments are increasingly available, yet dense language and distributed relationships among facts, evidence, reasoning, and rulings remain difficult for non-experts to interpret. Through a mixed-methods formative study with Chinese non-expert readers (survey N=34; interviews N=6), we identified structural, in...

Xin-Yi Chen, Rui-Ji Li, Yue-Lu Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and thre...

Rong Wang, Kun Sun, Ya-Dong Guo · 0 citations
#artificial intelligence Preprint Sep 2026

Before You Poll with LLMs: A Deliberative Diagnostic Framework

The Deliberative Polling Diagnostic Framework is introduced, which compares human and LLM belief shifts after identical informational interventions and offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.

A. Wali, Hassaan Tayyab · 0 citations
#artificial intelligence Open access Sep 2026

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

This work investigates an alternative standard designed to function despite ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for...

Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.