From scanner to scalpel: adapting IDEAL for accountable AI in radiology and image-guided neurosurgery
Abstract
Artificial intelligence (AI) is increasingly integrated into clinical decision-making in radiology and image-guided neurosurgery, yet evidence generation, oversight, and accountability remain uneven across the technology lifecycle. Strong performance on curated test sets may not persist across institutions, scanners, patient subgroups, or changing clinical workflows. Existing reporting guidance, medical-device standards, regulatory requirements, and machine learning operations (MLOps) address important parts of this problem, but are not designed to serve as a unified, stage-gated clinical-evidence pathway across the full lifecycle. We propose IDEAL-AI, an adaptation of the Idea, Development, Exploration, Assessment, and Long-term study principles for clinical AI. IDEAL-AI organizes evidence generation across five stages and three concurrent domains: validation; governance, regulation, and ethics; and patient and system outcomes. It is intended to complement, not replace, established reporting guidelines, regulatory requirements, facility-level quality programs, and MLOps infrastructure. By separating retrospective model development from prospective live exploration, requiring prespecified technical, fairness, human-factors, and clinical endpoints, and extending evaluation through post-market surveillance, the framework offers a practical way to determine when an AI system is ready to progress. Success should be judged not by adoption speed or benchmark accuracy alone, but by reproducible clinical benefit for patients, clinicians, and health systems.