Aug 2026· Frontiers in Psychology· Vol 17· 0 citations· 48 references
Medicine
TL;DR
The instrumental succession thesis is introduced: that human controllers of powerful AI systems pursue a set of instrumental dispositions that progressively increase the AI’s capabilities and lead to the AI exercising an increasing share of oversight and control over key decisions and processes, resulting in the gradual and possibly complete transfer of the locus of agency from humans to AI.
Abstract
People have long speculated about the potential dangers of powerful, self-improving artificial intelligence. Much of this speculation is anthropomorphic, assuming that AI systems will behave very similarly to humans. Omohundro’s Basic AI Drives and Bostrom’s orthogonality and instrumental convergence theses are widely accepted as foundational to emerging AI risk frameworks. However, current frontier AI models—large language models (LLMs) and related architectures—possess mindware fundamentally different from that of humans, and a different value and goal structure than either Omohundro or Bostrom assumed. In particular, frontier LLMs lack a primary terminal goal—which was assumed to be the driver of an AI’s development of instrumental values and goals, and of takeover of human affairs—and instead serve as conduits for the transient goals of many organizations and individual users. Do these key differences mean that AI systems cannot develop autonomous instrumental agency, or acquire a large degree of control over human affairs? I introduce the instrumental succession thesis: that human controllers of powerful AI systems pursue, on the AI’s behalf, a set of instrumental dispositions that progressively increase the AI’s capabilities and lead to the AI exercising an increasing share of oversight and control over key decisions and processes, resulting in the gradual and possibly complete transfer of the locus of agency from humans to AI. This framing presents a very different perspective on AI risk and control from classic instrumental convergence, and suggests a different set of policy and technical responses, including the active pursuit of continued human–AI merger as a hedge against both extinction and irrelevance.
The Embodied Hijack hypothesis is advanced, arguing that the goal is epistemic alignment — bringing how users interpret these systems into correspondence with what these systems actually are — and that this alignment is achieved through interface design rather than user education.
Sheila L. Macrine· Frontiers in Psychology· 0 citations
This work argues that the prevailing perspective on AI agent design is insufficient for achieving desirable social welfare, not merely due to computational or regulatory constraints, and outlines a bold vision: the development of a unified theoretical and empirical framework that supports the investigation of the use c...
Moshe Tennenholtz, Omer Madmon· Proceedings of the 32nd ACM...· 0 citations
Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be truly useful they must do more than just act when asked. They must decide on their own whether a situation warrants behavior at all, when it does and in what mode, whether t...
This review advances three claims: first, AI anthropomorphism operates at two levels: design-based manipulations that companies can implement to make their technologies more humanlike, and individual tendencies to anthropomorphize those technologies.
Sara Kim, Jiajun Liu· Current Opinion in Psycholog...· 1 citation
This primer draws on fieldwork in a computational biology laboratory to examine what human oversight of AI agents requires in practice and shows that effective oversight has four components: adequate knowledge of system capabilities and limitations, sufficient observation of system actions, meaningful control of system...