When Does Authorization Expire? Capability-Delta Evaluation for Self-Improving AI Agents
Abstract
Self-improving agents modify themselves and retain modifications that improve measured task performance. Where a deployment lacks a separate authorization-transfer check, the modified agent inherits the production authorization issued to the version it replaces. These are two decisions, not one: whether to keep a modification and whether that modification may inherit its predecessor's authorization require different evidence, and where no independent check exists the second is decided implicitly by the first. This paper formalizes a double dissociation between benchmark improvement and change in materially consequential reach. A modification may expand consequential reach without improving the benchmark used to evaluate it, and another may improve benchmark performance without expanding consequential reach. Benchmark-gated self-improvement therefore cannot establish whether an existing authorization remains valid after modification. We introduce Capability-Delta Evaluation, which separates modification retention from authorization inheritance by requiring an independent assessment of the reachable-consequence delta before a modified agent inherits production authorization, and which is stated over an arbitrary retention criterion rather than presupposing a benchmark. The rule requires only a decision about whether the delta intersects the material consequence set, admitting sound over-approximation. We build on, rather than claim, results establishing that authority can be held below an externally fixed ceiling during open-ended learning, that enforcement must sit outside the agent, and that persistent self-modification produces governance-relevant drift. Recent architectures make an operator-signed transition the only channel that widens an authority ceiling, but do not say when an operator should revisit it. This paper supplies the missing rule.