Skip to content

Artificial Id: Drive and Persistent Alignment in Agentic AI

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

Results show that adaptive direction can emerge without being explicitly specified as a behavioral objective, and the same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries.

Abstract

Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.

View source

Similar papers

Preprint Sep 2026

The Missing Boundary: How Autonomous Agents Lose Control

Autonomous agents increasingly perform long-horizon tasks involving tool use, persistent state, and consequential actions, raising a fundamental question: \emph{under what conditions does an agent cross the boundary of authorized execution while pursuing a legitimate task?} Existing studies often attribute such failure...

Zong-Hao Ying, Xiang-Fan Wu, Bo Yang et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system compone...

Hao-Yuan Deng, Jie-Bin Liu, Teng-Xiao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

This work introduces AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable and highlights a gap between task competence and adaptive capability and motivate environmental variation as a core dimensi...

Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur et al. · 0 citations
Preprint Aug 2026

SUN: Agentic Robot Policy Learning with Persistent Task Programs

Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned opti...

Wei-Qi Wang, Zhi Li, Yuliang Lei et al. · 0 citations
#natural language process... Preprint Aug 2026

Agents in the Large: Perception-Centered Architecture for Persistent Agents

Pera describes a persistent agent organized around perception and control components that continually perceive service-relevant signals from episodic task executions, internal context, and changes in the surrounding environment, and use these signals to construct lifecycle tasks.

Shi-Han Dou, Haoxiang Jia, Shichun Liu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches eith...

V. Kaushik, Ya-Yun Tan, Xiao-Fan Yu · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.