Sep 2026· The Paris Journal on AI & Digital Ethics· 1 citation· 60 references
Computer Science
TL;DR
Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, it is shown no model expresses a coherent policy across the three deployments, suggesting LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
Abstract
AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system’s behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation’s morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to 99 percentage points at a single escalation level, and a model’s success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
It is argued that demonstrating internal incoherence is a necessary precursor to AI alignment as well as a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs.
Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum· 0 citations
This paper argues that moral responsibility can fail under technologically mediated conditions in two structurally connected ways. First, agents may act from motivational states significantly shaped by external formative processes whose influence remains partially opaque to reflective awareness. Second, institutions ma...
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer e...
The article argues that many contemporary AI alignment practices risk a mistaken assimilation of moral agency to statistical learning. Techniques such as reinforcement learning from human feedback and constitutional AI often treat morality as a behavioral function that can be approximated from human discourse, behavior...
Evaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and...