Jul 2026· Journal of Clinical Medicine· Vol 15, pp. 5339· 0 citations· 19 references
Medicine
TL;DR
Machine learning models showed acceptable performance for predicting postoperative SSI after spinal surgery, suggesting that conventional statistical approaches may remain clinically useful in structured datasets.
Abstract
Background/Objectives: Surgical site infection (SSI) remains a clinically important complication after spinal surgery. This study developed and assessed machine learning approaches for predicting postoperative SSI using routinely collected preoperative clinical variables, with emphasis on calibration and clinical applicability. Methods: In this retrospective single-center study, four prediction models were developed in patients undergoing spinal surgery: logistic regression, random forest, gradient boosting, and XGBoost. Model training used five-fold stratified cross-validation, and performance was evaluated using a hold-out internal test set. Performance was assessed using the area under the receiver operating characteristic curve (AUC), area under the precision–recall curve (AUPRC), sensitivity, precision, F1 score, Brier score, and calibration slope. SHAP analysis was performed to evaluate model interpretability. Results: The incidence of SSI was 16.6%. In cross-validation, discrimination performance was broadly comparable across models, with logistic regression showing the highest observed AUC (0.814) and AUPRC (0.484). In the hold-out test set, the same model showed the highest AUC (AUC 0.806, 95% CI 0.757–0.852) and the highest sensitivity (0.758). Calibration performance varied across models. SHAP analysis identified C-reactive protein, hemoglobin, albumin, and white blood cell count as the most influential predictors. Perioperative variables provided only modest incremental predictive value. Conclusions: Machine learning models showed acceptable performance for predicting SSI after spinal surgery. Logistic regression demonstrated performance comparable to that of the evaluated machine learning models, suggesting that conventional statistical approaches may remain clinically useful in structured datasets. Preoperative clinical and laboratory variables were the major contributors to prediction, supporting their use for routine preoperative risk stratification.
Background Postoperative pulmonary complications (PPCs) are common adverse events after abdominal surgery in older adults, but existing risk scores may have limited transportability and clinical interpretability in elderly surgical populations. Methods This retrospective cohort study included 2,456 patients aged >=65 years who underwent abdominal surgery in the development/internal cohort and 542 patients in an independent external-validation cohort. Six algorithms were compared, including logistic regression, random forest, support vector machine, neural network, XGBoost, and LightGBM. Model performance was evaluated using discrimination, calibration, decision curve analysis, and external validation. SHAP analysis was used to support interpretability. Additional revision analyses examined pulmonary-function-test missingness, minor versus major PPCs, PPC co-occurrence patterns, temporal stability, COVID-era effects, and comparator-score performance. Results PPCs occurred in 425 of 2,456 patients (17.3%) in the development/internal cohort and 105 of 542 patients (19.4%) in the external-validation cohort. XGBoost showed the best overall performance, with AUCs of 0.856 (95% CI, 0.811–0.900) in the independent test set and 0.821 (95% CI, 0.781–0.861) in the external-validation cohort. At a 20% risk threshold, the independent-test sensitivity, specificity, PPV, and NPV were 87.5%, 60.9%, 32.0%, and 95.9%, respectively. SHAP analysis identified ASA physical status, COPD, upper abdominal surgery, age, emergency surgery, albumin, and surgical duration as leading contributors. The model separated patients into low-, moderate-, and high-risk groups with observed PPC rates of 6.9%, 24.4%, and 37.6%. Sensitivity analyses supported robustness to pulmonary-function missingness and temporal variation. Conclusion An interpretable gradient-boosting model may support risk-stratified perioperative assessment for elderly patients undergoing abdominal surgery. Prospective multicenter validation is required before routine clinical implementation.
Qiang Zhong, Guiming Huang, Wen Zhou et al.· Frontiers in Surgery· 0 citations
Surgical site infections (SSIs) are a common complication in gastrointestinal surgery, leading to major morbidity, mortality, and economic cost. There is a paucity of prediction models available for SSIs to improve the identification of patients at risk of an SSI. This review aims to evaluate the performance, validation, and methodological quality of prediction models for SSI in gastrointestinal surgery.
A systematic review was conducted of MEDLINE, Embase, and Web of Science databases from January 1, 2015, to July 3, 2025. The primary outcome was discriminative performance (area under the receiver operating characteristic curve [AUROC]). Secondary outcomes included calibration, clinical utility assessment, and validation.
From 7,692 records, 40 studies met the inclusion criteria, describing 129 distinct prediction models (86 regression-based and 37 machine learning/artificial intelligence-based). SSI incidence varied from 0.7% to 54.8%. AUROC for regression models ranged from 0.49 to 0.997 (median 0.76), and for ML/AI models from 0.50 to 0.991 (median 0.67). 27 models (20.93%) reported any form of calibration, and only 13 models (10.07%) showed a decision curve analysis.9 studies (47.50%) performed some form of external validation either of their new score and/or of a previous score, and 7 studies (17.5%) performed no validation of their newly developed score.
Contemporary SSI prediction models for gastrointestinal surgery remain characterised by inadequate validation, poor calibration reporting, insufficient assessment of clinical utility, and limited integration into electronic health records. These significant barriers must be addressed in future model development and validation to affect clinical practice.
H. Bhatti, S. Erridge, Artemis Mantzavinou et al.· British Journal of Surgery· 0 citations
Background: Cases requiring 13 or more tissue sections in Mohs micrographic surgery (MMS) demand extended operative time, additional resources, and often specialised closure techniques. Pre-operative identification of such cases would improve surgical scheduling, resource allocation, and patient counselling. We aimed to develop and validate a machine learning prediction tool using pre-operative clinical features to identify cases likely to require13 sections. Objectives: To develop and validate machine learning models for predicting which Mohs procedures will require 13 sections, using pre-operative clinical features, and to identify key predictive factors. Methods: We analysed 408 consecutive Mohs procedures with 16 pre-operative clinical variables. Thirty machine learning algorithms were evaluated, including ensemble methods (Stacking, Voting), gradient boosting (XGBoost, LightGBM, CatBoost), neural networks (3-7 layers), support vector machines, and traditional classifiers. Model performance was assessed using 5-fold stratified cross-validation and independent test set evaluation. Feature importance was determined using SHAP (SHapley Additive exPlanations) analysis. Results: The stacking ensemble achieved the highest cross-validation AUC of 0.891 (95% CI: 0.849-0.934) and test AUC of 0.884. Tumour area (cm2), calculated using the ellipse formula to approximate clinical tumour morphology, emerged as the strongest predictor (SHAP importance: 0.141), followed by tumour size dimensions (0.086 and 0.068), aggressive histopathology (0.046), and recurrence status (0.035). Wide neural network architectures (5-layer) outperformed deeper configurations (7-layer). The model demonstrated 70.7% high-confidence predictions with uncertainty <15%. Conclusions: Machine learning models using pre-operative clinical features can accurately predict which Mohs procedures will require 13 or more sections. The stacking ensemble approach provides robust predictions suitable for clinical decision support. External validation in multi-centre cohorts with diverse patient populations and practice patterns is warranted to assess model generalisability.
Y. A. Aksoy, S. Lee, G. Moreno-Bonilla· medRxiv· 0 citations
Patients with lumbar spinal stenosis are typically elderly with multiple comorbidities, necessitating accurate preoperative anesthetic risk assessment. The American Society of Anesthesiologists (ASA) classification quantifies functional reserve and disease burden, serving as a widely used tool for risk stratification. However, ASA classification is often influenced by subjective factors including physician experience and varies among clinicians with different seniority, while the assessment process remains time-consuming. This study aimed to develop an automated model for anesthetic risk stratification and evaluate its performance, with the goal of providing decision support for surgical and anesthetic management in this patient population.
Clinical data of 600 patients with lumbar spinal stenosis were collected and randomly divided into training (
n
= 480) and internal validation (
n
= 120) sets. An additional 100 patients from another tertiary hospital formed an external validation set. The model was validated and hyperparameter-tuned using k-fold cross-validation. Model performance, including overall classification and high-risk identification, was evaluated using accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa, linear weighted accuracy, positive predictive value, and negative predictive value.
In the internal validation set, the accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa coefficient and linear weighted accuracy of the model are 0.97, 0.96, 0.95, 0.96, 0.93 and 0.98 respectively, while in the external validation set, they are 0.97, 0.97, 0.94, 0.96, 0.93 and 0.99 respectively. The confusion matrix heatmap shows that the error is mainly concentrated between adjacent classes, and there is no cross-class misjudgment. In the internal validation set, the model's accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and Kappa coefficient for identifying high-risk patients were 0.98, 0.95, 0.98, 0.91, 0.99, and 0.92, respectively. In the external validation set, these values were 0.98, 0.93, 0.99, 0.93, 0.99, and 0.92, respectively.
The machine learning model developed in this study demonstrates strong capability in stratifying anesthetic risk for patients with lumbar spinal stenosis, providing valuable reference for selecting surgical and anesthetic approaches.
Jitao Yang, Yixi Wang, Qihao Chen et al.· Frontiers in Medicine· 0 citations
Postoperative wound healing complications present a major challenge in plastic and reconstructive surgery, prolonging recovery and impairing outcomes. Early risk identification is difficult due to complex interactions among clinical, laboratory, and molecular factors. This study developed and evaluated machine-learning (ML) models to predict wound healing outcomes and identify key complication predictors. Utilizing a dataset of 95 women and 76 variables (including hematological, biochemical, coagulation, and gene expression profiles), we evaluated several ML approaches, including Decision Tree, Extra Trees, Gaussian/Bernoulli Naive Bayes, Logistic Regression, and Support Vector Machine. Model performance was assessed via k-fold cross-validation, ROC analysis, and SHAP feature importance. Molecular markers (COL1A1, MMP9, MAPK1, MAPK8, IL10, and CCL2) emerged as the strongest predictors, whereas conventional clinical variables showed limited value. The models achieved high discriminative performance, with validation ROC–AUC values ranging from 0.903 to 0.913. Extra Trees and Gaussian Naive Bayes demonstrated the highest sensitivity for detecting complications (Recall = 0.820 ± 0.238 and 0.807 ± 0.246, respectively). These findings highlight the value of integrating molecular-genetic biomarkers with ML for personalized risk stratification and preventive care in reconstructive surgery.
L. Sydorchuk, R. Gumennyi, Miroslav Škoda et al.· Computation· 1 citation