From AI Safety via Debate to Evidence-Grounded Adversarial Assurance
Abstract
Can competing AI systems help humans evaluate reasoning they cannot independently verify? AI safety via debate proposes this possibility. Yet winning an argument is not a certificate of safe action. This paper develops evidence-grounded adversarial assurance: a mathematical framework that separates proposal generation, adversarial scrutiny, independent evidence verification, and execution authority. Its central contribution is a compositional calculus linking local safety certificates to system-level guarantees while distinguishing operational risk from statistical confidence. Drawing on category theory, probabilistic program semantics, abstract interpretation, and control theory, the framework specifies when assurance survives sequential composition, component replacement, adaptive certificate selection, and recurrent operation. The analysis addresses bounded incentives, omitted hazards, correlated reviewers, distribution shift, and unreliable safety labels. Conditional theorems, counterexamples, numerical illustrations, and an executable certificate checker expose the proposed architecture's capabilities and limitations. A preregistrable evaluation program targets software review, research auditing, tool authorization, and stochastic decision support, measuring safety alongside useful task completion. The contribution is not a claim that debate solves alignment or that deployment safety has been demonstrated. It is a rigorous, testable foundation for determining when adversarial reasoning, independent verification, and constrained execution can support defensible safety guarantees under explicitly stated assumptions.