From AI Safety via Debate to Evidence-Grounded Adversarial Assurance
Can competing AI systems help humans evaluate reasoning they cannot independently verify? AI safety via debate proposes this possibility. Yet winning an argument is not a certificate of safe action. This paper develops evidence-grounded adversarial assurance: a mathematical framework that separates proposal generation,...