Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis
Abstract
To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes. AI models mean performance match or exceed radiologist performance in neuroblastoma diagnosis. MLM outperform HNs in differential diagnosis. AI nomograms and gene signatures showed higher descriptive AUCs than several conventional prognostic markers, but formal comparative inference was not possible. Only 33.9% of models were calibrated; 24.5% underwent external validation. AI must transition from proof-of-concept to prospective clinical validation. AI models mean performance match or exceed radiologist performance in neuroblastoma diagnosis. MLM outperform HNs in differential diagnosis. AI nomograms and gene signatures showed higher descriptive AUCs than several conventional prognostic markers, but formal comparative inference was not possible. Only 33.9% of models were calibrated; 24.5% underwent external validation. AI must transition from proof-of-concept to prospective clinical validation.