Sep 2026· The Paris Journal on AI & Digital Ethics· 1 citation· 50 references
Computer Science
TL;DR
This work investigates an alternative standard designed to function despite ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for argument cogency.
Abstract
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton’s theory of argumentation schemes and Govier’s criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items—6,778 judge-scored cells, validated against 89.6% inter-judge agreement on the binary failure judgment—models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas (≥ 20% per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer e...
Large Language Models (LLMs) are entering democratic contexts as instruments of governance, where the challenges at hand are ill-structured, marked by ambiguity and contestation. Ill-structured democratic problems demand more than factual precision; they call for intersubjective reasoning: context-sensitive judgments t...
Francesco Veri, Gustavo Kreia Umbelino· Proceedings of the National...· 0 citations
It is argued that once moral trustworthiness is set aside, the central normative question shifts from whether AI systems can be trusted to how their epistemic influence on human decision-making ought to be justified and governed.
John Dorsch, Maximilian Moll, Ophélia Deroy· Philosophy & Technology· 0 citations
This critical integrative review synthesizes research on automation bias, algorithm appreciation and aversion, large-language-model sycophancy, persuasive dialogue, anthropomorphic interaction, explainable AI, and human autonomy to reposition productive disagreement as a positive safety and autonomy property of human–A...
Modern states increasingly use AI systems in public decision-making, including benefits eligibility, risk classification, immigration adjudication, and judicial administration. Defenders often argue that democratic authorization, validation, and review are sufficient to render such systems legitimate; critics often rep...
Sai-Ming Wong· Ethics and Information Techn...· 0 citations