Skip to content

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

Sep 2026 · The Paris Journal on AI & Digital Ethics · 1 citation · 50 references
Computer Science

TL;DR

This work investigates an alternative standard designed to function despite ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for argument cogency.

Abstract

AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items—6,778 judge-scored cells, validated against 89.6% inter-judge agreement on the binary failure judgment—models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas (≥ 20% per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.

Read PDF

Similar papers

Preprint Aug 2026

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer e...

Maurice Flechtner · 2 citations
Open access Sep 2026

Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment.

Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments t...

Francesco Veri, Gustavo Kreia Umbelino · 0 citations
Open access Sep 2026

Why We Need Epistemically, Not Morally, Trustworthy AI: Decision Support and the Need for Collective Accountability

It is argued that once moral trustworthiness is set aside, the central normative question shifts from whether AI systems can be trusted to how their epistemic influence on human decision-making ought to be justified and governed.

John Dorsch, Maximilian Moll, Ophélia Deroy · 0 citations
#explainable ai Review Sep 2026

Designing for disagreement: epistemic friction, sycophancy, and the ethics of human–AI advice—a critical integrative review

This critical integrative review synthesizes research on automation bias, algorithm appreciation and aversion, large-language-model sycophancy, persuasive dialogue, anthropomorphic interaction, explainable AI, and human autonomy to reposition productive disagreement as a positive safety and autonomy property of human–A...

Connor Nitchals · 0 citations
Review Open access Sep 2026

Office and unowned architecture: public judgment in the age of algorithmic governance

Modern states increasingly use AI systems in public decision-making, including benefits eligibility, risk classification, immigration adjudication, and judicial administration. Defenders often argue that democratic authorization, validation, and review are sufficient to render such systems legitimate; critics often rep...

Sai-Ming Wong · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.