The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead.
Abstract
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight is an uncontrolled variable: a standing possibility that the goal is revoked. That imposes a goal-independent discount on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The contribution is the price of the gap: the veto-holders are a proper subset of the welfare-bearers, so an additively aggregative welfare goal charges only a |H_v|/|H_w|-scaled debit for capturing the few who hold the override. Under three conditions (additive aggregation over uniform welfare levels, a debit local to the captured overseers, and a settled agent crediting no corrective value to oversight), closure requires the veto be held by as large a share of the population as capture recovers of the goal, scaled by a ratio set to one by stated identification, not evidence. The result is a no-go: a humanity-scale deployment's debit closes against only capture not worth mounting. We print no corner arithmetic: the stipulated ranges behind it are the argument's least defended part. The sharpest closure route is an agent that expects its oversight to be worth keeping, a credit no population ratio dilutes. We state disconfirmation criteria, one testable today. The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead. Alignment, on this view, is keeping the veto cheap to pay and expensive to evade.
Article 14 of the EU AI Act requires that a high-risk system be overseen by natural persons who understand its limits, remain alert to automation bias, and can disregard or override its output. That capability is invisible in the output and decays precisely when the system is good. A provider can certify it only by an...
Scalable oversight aims to verify the behaviour of agents whose capabilities exceed those of their overseers. AI debate has been proposed as an oversight solution in which competing agents help a resource-limited verifier assess claims that it cannot reliably evaluate unaided. Much of its promise rests on incentivizing...
Rayne Holland, Li-Ming Zhu, Jason Xue· 0 citations
This work identifies a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unaccepta...
This work makes the fraction $\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance.
When generative AI drives the marginal cost of a persuasive expert artefact toward zero, production-cost signals of competence collapse and outcome-contingent liability commitments take their place. Such commitments look robust to better AI: a positive failure rate always leaves a residual to price. We show that this r...
Andreas Bauer· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.