Artificial intelligence and computational prediction models for risk stratification, treatment response, and outcomes in colorectal cancer: a narrative review
Abstract
To critically evaluate, within a decision-centred framework, the clinical readiness of artificial-intelligence (AI) and computational prediction models used for risk stratification, treatment-response assessment, and outcome prediction in colorectal cancer (CRC), while distinguishing tumour site, intended clinical decision, assessment time point, and validation level, and identifying the gap between predictive performance and clinical impact. PubMed/MEDLINE and Web of Science were searched for English-language studies published from January 2010 through 31 May 2026; ScienceDirect, Google Scholar, reference lists, and forward citation tracking were used as supplementary sources. Two reviewers independently screened studies and extracted design, endpoint, data modality, unit of analysis, validation, calibration, decision analysis, workflow evaluation, accessibility, and clinical-impact evidence. Study-level methodological features were assessed separately from task-level translational readiness. Comparatively stronger evidence was found for computer-aided endoscopic assessment of invasion depth, MRI-based lymph-node prediction in rectal cancer, H&E whole-slide-image pre-screening for MSI/MMR status, and selected externally tested models for pathological complete response. Evidence remained early-stage for CT-based nodal staging in colon cancer, tumour mutational burden prediction, neoadjuvant immunotherapy-response prediction, synchronous distant-metastasis assessment, recurrence, survival, and most perioperative-risk models. Prospective validation was uncommon and was not equivalent to demonstration of clinical benefit. A decision-centred synthesis indicates that CRC AI is most mature as site-specific and task-specific reader assistance or triage rather than autonomous decision-making. Clinical translation requires representative external validation, calibration, predefined action thresholds, model availability, workflow testing, safety monitoring, and prospective evidence that model-guided care improves decisions or outcomes.