Skip to content
Open access

Preoperative artificial intelligence-based risk model for surgical reintervention after microsurgical free flap reconstruction.

Jul 2026 · Journal of Plastic, Reconstructive & Aesthetic Surgery · Vol 120, pp. 290-299 · 0 citations · 25 references
Medicine

TL;DR

XGBoost is retained as the principal model based on combined superiority in discrimination and calibration, with random forest as a robust comparator, and prospective external validation with recalibration is required before clinical adoption.

Abstract

Background

Flap-related vascular complications requiring surgical reintervention remain a source of morbidity after microsurgical free flap reconstruction and preoperative risk estimation relies on clinical judgement. Therefore, we developed and internally validated a strictly preoperative multivariable model for this outcome.

Methods

A retrospective cohort of 650 consecutive adults at two high-volume referral centres (649 analysable) was analysed. The primary outcome was unplanned re-exploration within 30 days for arterial, venous, mixed thrombosis, or clinically significant vasospasm. Nineteen preoperative predictors were considered; intraoperative and surgeon variables were excluded. Least absolute shrinkage and selection operator (LASSO), random forest, and XGBoost were fitted on a 70% training partition and evaluated on a 30% test set. Performance was assessed using discrimination (AUC), calibration (intercept, slope, ICI, and E/O), and Brier score. Reporting followed TRIPOD; risk of bias used the four-domain PROBAST.

Results

Event rate was 17.1% (111/649). XGBoost achieved the highest discrimination (AUC 0.84; 95% CI 0.74-0.92) and relatively better calibration (intercept 0.02, slope 0.87, ICI 0.039, E/O 0.89). Random forest showed comparable discrimination (AUC 0.82; 0.71-0.90) but poorer calibration (slope 1.33). LASSO demonstrated the lowest discrimination (AUC 0.79; 0.69-0.88). Prior oncologic history, surgical indication, and flap composition were influential predictors. A decile table from XGBoost showed a monotonic gradient, with events rising from ≤10% in lower deciles to 80% (95% CI 58-92) in the top decile.

Conclusions

XGBoost is retained as the principal model based on combined superiority in discrimination and calibration, with random forest as a robust comparator. The model is not decision-ready, and prospective external validation with recalibration is required before clinical adoption. LAY SUMMARY Using 649 analysable free flap reconstructions, we developed preoperative models to predict unplanned surgical reintervention for flap-related vascular complications within 30 days. XGBoost performed the best (AUC 0.84), supporting risk stratification before surgery, but prospective external validation is needed before clinical use.

Read PDF

Similar papers

Aug 2026

Operative-Time Gradients Reveal Miscalibration in Microsurgical Risk Prediction.

BACKGROUND General surgical risk models support perioperative counselling, but acceptable overall calibration may conceal clinically important error. We evaluated generalized American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) predictions after microsurgical reconstruction and tested whether morbidity calibration varied across operative time. METHODS Adult ACS NSQIP cases from 2014-2023 were analyzed. Plastic-surgery-service cases meeting exact principal Current Procedural Terminology (CPT) or principal-procedure-text criteria formed the primary cohort. Performance was assessed using observed-to-expected (O:E) ratios, discrimination, and calibration measures. Operative-time calibration was examined by quartiles, deciles, and continuous spline models. Sensitivity analyses used exact principal CPT codes alone and excluded cases with a known morbidity component or reoperation on postoperative day 0 or 1. RESULTS Among 20,604 microsurgery cases, observed and predicted morbidity were similar overall (12.52% vs. 12.79%; O:E, 0.98; 95% confidence interval [CI], 0.94-1.02), although discrimination was modest (area under the receiver operating characteristic curve, 0.639; 95% CI, 0.627-0.651). Mortality was uncommon (21 deaths; O:E, 0.81; 95% CI, 0.50-1.24). Aggregate calibration concealed a graded reversal across operative time. Morbidity was overpredicted in the shortest quartile (8.66% observed vs. 12.87% predicted; O:E, 0.67; 95% CI, 0.61-0.74) and underpredicted in the longest quartile (16.94% vs. 12.96%; O:E, 1.31; 95% CI, 1.22-1.40). The prespecified linear interaction was nonsignificant (P = 0.183), whereas flexible continuous analyses demonstrated calibration variation before and after global recalibration (both P < 0.001). Findings persisted in both sensitivity analyses. CONCLUSIONS Generalized ACS NSQIP morbidity prediction was accurate in aggregate but systematically miscalibrated across the operative-time gradient, overestimating risk in shorter operations and underestimating it in longer operations. Realized operative time should be interpreted as a postoperative marker of incompletely captured procedural complexity and intraoperative course, not as a causal exposure or preoperative predictor.

Shaan Sekhon, Miracle Uzoekwe, C. Tompkins-Rhoades et al. · 0 citations
Jul 2026

AI Risk Prediction Tools for Autologous Breast Reconstruction.

BACKGROUND Autologous breast reconstruction offers patients a durable and natural-appearing option after mastectomy. However, complication risks include flap loss, infection, and delayed wound healing. This study developed both traditional statistical and machine learning (ML) models to predict the risk of developing a 90-day postoperative complication after autologous reconstruction. METHODS Patient data were retrospectively collected for patients who underwent abdominal-based autologous breast reconstruction at Memorial Sloan Kettering Cancer Center (January 2015-September 2024). Multivariable logistic regression models and supervised ML models were developed to predict infection, hematoma, seroma, delayed wound healing, and flap compromise. RESULTS 2,128 patients (3,249 flap reconstructions) were included. Overall, 90-day complications occurred in 475 (22.3%) patients, including infection (10%), hematoma (6%), delayed healing (4.1%), seroma (3.9%), and flap compromise (2.8%). AUCs ranged from 0.60 to 0.66 (logistic regression) and 0.64 to 0.73 (ML). Higher BMI was associated with increased risk for seroma (OR: 1.1, 95% CI: 1.04-1.13) and neoadjuvant chemotherapy for infection (OR:1.72, 95% CI: 1.17-2.51). Key ML model predictors on SHAP analysis included age, BMI, and pre-reconstruction radiation. CONCLUSION Individualized risk prediction models for complications after autologous breast reconstruction were developed using traditional statistics and ML. These models may help personalize treatment; however, further research is needed to enhance model predictive performance and clinical utility.

Jonlin Chen, Ariel Gabay, Abbas M. Hassan et al. · 0 citations
May 2026

Surgical Decision-Making in Anterior and Middle Skull Base Lesions: A Single-Center Experience with Comparative Outcomes and a Preoperative Predictive Model

Abstract Introduction Surgical management of anterior and middle skull base (ASB/MSB) lesions requires careful balancing of maximal resection and preservation of neurological function. Although both transcranial (TC) and endoscopic endonasal (EE) approaches are widely used, the extent to which the surgical approach independently influences outcomes remains unclear. Methods We conducted a retrospective, single-center, observational study of adult patients who underwent surgery for ASB/MSB lesions. Demographic, clinical, radiological, and surgical variables were analyzed. Outcomes included extent of resection (EOR), complications, intraoperative blood loss, operative time, and length of hospital stay. Comparative analyses between TC and EE approaches were performed and interpreted as exploratory. A preoperative predictive model for gross total resection (GTR) was developed using Least Absolute Shrinkage and Selection Operator-penalized logistic regression with internal bootstrap validation. Results A total of 84 patients were included. TC approaches were performed in 62 patients (73.8%) and EE in 22 (26.2%). GTR was achieved in 77.4% of cases, with no significant difference between approaches. The EE group showed significantly lower intraoperative blood loss, shorter operative time, and reduced length of hospital stay. Complication rates did not differ significantly between groups. The predictive model demonstrated good apparent discrimination (area under the curve 0.97). Conclusions Our findings suggest that tumor characteristics may play a substantial role in determining outcomes; however, the observational design does not allow independent assessment of the causal effect of the surgical approach. When appropriately selected, both TC and EE strategies achieve comparable effectiveness and safety. At this stage, the proposed model may serve only as an exploratory predictive tool.

L. Bonosi, U.E. Benigno, G. Giammalva et al. · 0 citations
Open access Aug 2026

A LASSO regression-driven nomogram to predict functional outcomes after surgery for chronic lateral ankle instability

Objective This study aimed to identify predictors of postoperative recovery in chronic lateral ankle instability (CLAI) and construct a risk-estimation nomogram for individualized risk assessment. Methods Between January 2021 and December 2023, 132 patients with CLAI who underwent arthroscopic Broström repair with suture-tape augmentation were retrospectively included and classified into good-recovery [American Orthopaedic Foot and Ankle Society (AOFAS) ankle-hindfoot score ≥75, n = 73] and poor-recovery groups (AOFAS <75, n = 59). Candidate predictors were identified using univariate analysis, LASSO regression, and multivariable logistic regression. A nomogram was developed and internally validated using 1,000 bootstrap resamples, and model performance was assessed by ROC curve analysis, calibration analysis, and decision curve analysis (DCA). Results Six variables were identified as independently associated with poor postoperative recovery: sprain count (OR = 1.439, P = 0.043), osteophytes (OR = 5.951, P = 0.004), syndesmosis injury (OR = 5.272, P = 0.007), Outerbridge grade II cartilage damage (OR = 23.928, P = 0.017), thin or absent ATFL remnant <1.0 mm (OR = 5.256, P = 0.029), and CFL/ATFL angle <70° (OR = 11.598, P = 0.003). The nomogram showed acceptable discrimination, with an apparent AUC of 0.907 and an optimism-corrected AUC of 0.876 after 1,000 bootstrap resamples. Calibration demonstrated acceptable agreement between predicted and observed outcomes, and DCA suggested potential clinical utility. Conclusion This study identified six variables associated with AOFAS-defined poor postoperative functional recovery in patients with CLAI and developed an internally validated nomogram for individualized risk estimation.

Feng Lin, Zhiyao Lv, Zhidong Zhao et al. · 0 citations
Open access Aug 2026

Machine learning–based prediction of major amputation risk after initial limb-preserving surgery in diabetic foot

Background Accurate preoperative prediction of whether an initially limb-preserving strategy in diabetic foot management will culminate in minor or major amputation remains a clinical challenge. This study aimed to develop and evaluate using two temporally separated cohorts a machine-learning framework using routinely available baseline clinical, laboratory, and selected imaging and vascular variables. Methods Two temporally separated cohorts were used, with Dataset 1 for model development and Dataset 2 for temporally separated evaluation. A 20-repetition stratified outer-split workflow was implemented, incorporating two-step feature selection, Optuna-based hyperparameter optimization, training-only SMOTE, and threshold tuning to maximize the F2-score under a recall constraint of ≥0.70. Six classifiers were evaluated using average precision (AP), ROC-AUC, recall, precision, specificity, accuracy, and Brier score. Results The major-amputation group exhibited a more severe baseline phenotype, including higher inflammatory burden, worse neuropathy and wound severity, and a higher prevalence of necrotizing fasciitis. Internally, multilayer perceptron achieved the highest AP (55.9% ± 13.6%). In external evaluation, k-nearest neighbors achieved the highest AP (65.1% ± 10.2%) and recall (72.8% ± 19.6%), whereas multilayer perceptron showed higher precision and specificity. Key contributors included necrotizing fasciitis, neuropathy severity, hemoglobin, PEDIS classification, and inflammatory indices. Conclusion These findings suggest that prediction of amputation level is feasible, validated in a temporally separated cohort, and clinically interpretable, and may support future decision-support applications, although further validation is required before clinical implementation.

Kerim Bora Yılmaz, Gokalp Tulum, F. Cuce et al. · 0 citations