Skip to content

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

Sep 2026 · The Paris Journal on AI & Digital Ethics · 1 citation · 60 references
Computer Science

TL;DR

Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, it is shown no model expresses a coherent policy across the three deployments, suggesting LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

Abstract

AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system’s behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation’s morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to 99 percentage points at a single escalation level, and a model’s success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

Read PDF

Similar papers

Preprint Aug 2026

Incoherent by Design? On the Moral Self-Consistency of LLMs

It is argued that demonstrating internal incoherence is a necessary precursor to AI alignment as well as a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs.

Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum · 0 citations
Open access Sep 2026

When responsibility cannot arise: algorithmic mediation and the preconditions of moral imputability

This paper argues that moral responsibility can fail under technologically mediated conditions in two structurally connected ways. First, agents may act from motivational states significantly shaped by external formative processes whose influence remains partially opaque to reflective awareness. Second, institutions ma...

Åke Elden · 0 citations
Preprint Aug 2026

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer e...

Maurice Flechtner · 2 citations
#generative ai Open access Sep 2026

Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner

The article argues that many contemporary AI alignment practices risk a mistaken assimilation of moral agency to statistical learning. Techniques such as reinforcement learning from human feedback and constitutional AI often treat morality as a behavioral function that can be approximated from human discourse, behavior...

Saša Josifović · 0 citations
#small language model Open access Sep 2026

Measuring structural value alignment in sixteen models: LLMs have human-like moral spaces

Evaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and...

Ye-Xiang Tang · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.