The evolutionary origin of value in biological organisms is traced by tracing the evolutionary origin of value in biological organisms to conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.
Abstract
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of"vicarious selectors"that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the"orthogonality thesis"separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.
The instrumental succession thesis is introduced: that human controllers of powerful AI systems pursue a set of instrumental dispositions that progressively increase the AI’s capabilities and lead to the AI exercising an increasing share of oversight and control over key decisions and processes, resulting in the gradua...
Preston W. Estep· Frontiers in Psychology· 0 citations
A unified framework connecting empirical hallmarks of consciousness attribution to a structured risk taxonomy of Seemingly Conscious AI (SCAI), AI systems that exhibit hallmarks which elicit consciousness attribution from users is provided.
Ben Bariach, P. Schoenegger, M. Bhaskar et al.· AI and Ethics· 3 citations
It is argued that AI systems themselves will increasingly participate in the reconstruction of the authors' shared epistemic environment because they readily supply narrative material and personalised interpretive scaffolding at precisely the moment when users'conceptual assumptions may already be loosened.
T. Pollak, H. Morrin, Murray Shanahan· 0 citations
The Embodied Hijack hypothesis is advanced, arguing that the goal is epistemic alignment — bringing how users interpret these systems into correspondence with what these systems actually are — and that this alignment is achieved through interface design rather than user education.
Sheila L. Macrine· Frontiers in Psychology· 0 citations
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as t...
Andrew Smart, Shazeda Ahmed, Jackie Kay et al.· 0 citations
Whether analytic theology can supply a symbolic logical framework apt to encode moral safeguards inside LLM pipelines is investigated, which targets the engineering of moral computation: a formal language placed between metaphysical content and machine implementation.
Fernando Negrini· International Journal of Bio...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.