Skip to content
Open access

On AI operational alignment and misalignment for high-risk AI systems: case study

Aug 2026 · Scientific Reports · 0 citations

TL;DR

The paper describes an approach to automated dynamic operational alignment in a complex high-risk process, exemplified by a case study on autonomous drilling, and possible sources of AI misalignments are discussed.

Abstract

AI alignment is generally associated with ethics and social aspects. For high-risk AI systems, principles such as safety, stability and performance are crucial for their adoption in real-world application, to build trust and increase efficiency of a process. By safety it is meant both safety of the system itself but also safety of the process and environment. Any decision taken by a high-risk AI system should preserve safety, ensure continuous operation, real-time functioning, smooth control, and improve process efficiency. Human oversight needs to be continuously ensured, both to preserve safety of the process in case of transition from autonomous to manual mode in case of failures, but also to allow for contextual knowledge of the decisions. All these objectives are included in the AI operational alignment. High-risk AI systems designed to automatize complex processes in critical environments are continuously adapting to dynamic contexts, thus a static alignment might not be sufficient to properly assess their behaviour. The paper describes an approach to automated dynamic operational alignment in a complex high-risk process, exemplified by a case study on autonomous drilling. Possible sources of AI misalignments in this case are discussed and their potential implications.

Read PDF

Similar papers

2026

Human Authority in the Age of Intelligent Systems: A Sociotechnical Framework for Decision Assurance and Psychosocial Safety

The rapid integration of intelligent systems into high-risk industries is reshaping how work is performed, decisions are made, and responsibility is distributed. While automation and artificial intelligence promise gains in efficiency, safety, and productivity, they also introduce under-recognised sociotechnical risks, including degraded human decision authority, cognitive overload, loss of practical expertise, and increasing psychosocial strain associated with uncertainty and workforce transition. This paper argues that current approaches to human factors engineering and system design are insufficient to address these conditions. While technical assurance frameworks for AI-enabled systems continue to mature, equivalent structures for ensuring the integrity, clarity, and sustainability of human decision-making remain underdeveloped. This gap presents a critical risk in high-reliability environments where human judgment is essential for managing variability, uncertainty, and system failure. Drawing on applied experience in automation-intensive operations and informed by an emerging global research agenda, this research introduces the Human Authority Index (HAI) as a novel sociotechnical framework developed by the author for assessing and designing decision-making in human–machine systems. The HAI positions human authority as a safety-critical variable across three dimensions: Decision Clarity, Cognitive Sustainability, and Psychosocial Safety. Together, these dimensions extend traditional human factors approaches by integrating organisational psychology, sociotechnical systems theory, and resilience engineering into a unified model. Through conceptual analysis and industry-informed scenarios, the paper demonstrates how poorly integrated automation can create hidden system fragility, while systems designed with clear human authority exhibit greater resilience and workforce engagement. This research contributes a practical framework to support safe, effective, and meaningful human participation in increasingly autonomous environments.

Peta Chirgwin · 0 citations
Conference Aug 2026

Toward Operationalizing Pluralistic Normative Principles for AI Systems

The ethical issues surrounding the development and deployment of AI are becoming very critical as these technologies evolve and affect more facets of society. To make AI systems more beneficial and reduce their potential negative impact on public life, international communities, as well as local and global organizations, are working to elicit normative principles that prescribe what AI systems should and should not do in order to remain trustworthy. However, these principles are new, ambiguous, and very high-level, and there are still no clear standards guiding developers on how to operationalize them in practice. I argue that, operationalizing such principles requires collaboration with multidisciplinary experts to ground them in the context of specific AI solutions and their deployment environments. Otherwise, the problem remains open-ended and difficult to apply in practice. In my thesis, I aim to address this gap by building a bridge between multidisciplinary normative experts and AI developers to operationalize high-level normative principles for AI systems with open capabilities but specific tasks and deployment contexts. I present here my early results, current progress, and future research directions.

L. Abodinar · 0 citations
Review

Trustworthy AI for Safety-Critical Perception and Decision Systems

It is asserted that trustworthiness is a systems property, not a single algorithmic feature, and that realising it requires coordinated advances in explainability, uncertainty quantification, human–AI interaction design, failure detection, and domain-specific data governance.

Shruti Kshirsagar · 0 citations
#artificial intelligence Preprint Sep 2026

AI Deployment Accountability Engineering: A Vision for Accountable AI in Safety-Critical Socio-Technical Systems

Artificial intelligence systems are rapidly becoming critical components in healthcare, finance, public services, and other safety-critical domains. Yet the engineering practices used to evaluate these systems remain predominantly model-centric, emphasizing properties such as accuracy, robustness, fairness, and interpretability before deployment. These properties are necessary but insufficient once an AI system operates within an ever changing socio-technical environment characterized by distribution shifts, institutional constraints, human feedback loops, privacy requirements, and interactions among multiple AI agents. This vision paper introduces AI Deployment Accountability Engineering (ADAE), a proposed AI engineering subdiscipline concerned with establishing measurable, continuous, and actionable accountability for deployed AI systems. ADAE treats accountability as a deployment-layer property rather than solely as a property of an individual model. It seeks to determine whether an AI-enabled system continues to operate within acceptable risk limits, identify the contexts in which failures emerge, attribute failures across interacting technical and human components, translate technical failures into downstream consequences, and support timely intervention. We articulate a research agenda built around four interconnected pillars: structured discovery of context-dependent failure modes, privacy-preserving accountability measurement, system-level risk analysis for agentic AI, and translation of technical failures into operational, and institutional risks. The broader goal is to establish foundational principles, mathematical tools, and system architectures for accountable AI deployment across safety-critical applications.

Murat Kantarcioglu · 0 citations
Open access Sep 2026

Wild, Thick, and Wicked: Situated Evidence on AI-In-Use for Decisions About Deploying AI Systems

Organizations are adopting generative AI faster than the evidence base needed to govern it. Existing evaluation tools such as benchmarks, alignment scores, and safety tests were built for model development, not for judging whether systems will create value, introduce friction, or shift risk in specific real-world settings. As a result, there is little systematic evidence about how AI behaves once it is embedded in everyday work. This paper proposes a real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time. Instead of treating variability across users, tasks, and settings as noise to be controlled away, the framework treats that variation as the central source of deployment-relevant evidence. It sets out four design principles for producing decision-ready evidence at scale and proposes a shared evaluation architecture combining a structured observation environment, a metrics hub, and reusable consortium models that summarize system behavior across contexts. Rather than replacing traditional benchmarks, this framework adds a sociotechnical evidence layer that connects model capabilities to the organizational and practitioner level outcomes where deployment decisions are actually made.

Reva Schwartz, Gabriella Waters · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.