A Repeated-Evaluation Comparison of Traditional, Machine Learning, and Deep Learning Survival Models Across Static and Longitudinal Data
Survival analysis plays a central role in medical research. Although the Cox proportional hazards (CoxPH) model remains the standard approach, machine learning and deep learning methods have been increasingly adopted. However, many published comparisons have relied on a single train–test split, which may produce unreliable performance estimates, particularly for unstable modelling approaches. This study compared CoxPH (LASSO-selected), Random Survival Forest (RSF), and Long Short-Term Memory (LSTM) networks using four survival datasets: breast cancer (n=4024), heart failure (n=299), recidivism (n=4618), and the Mayo Clinic Primary Biliary Cholangitis Sequential Dataset (PBC2; n=312) containing time-varying covariates. Model performance was evaluated using a two-stage protocol comprising a conventional 80–20 stratified train–test split followed by 100 repeated stratified 80–20 train–test splits. Performance was assessed using the concordance index (C-index), integrated Brier score (IBS), and time-dependent area under the receiver operating characteristic curve (AUC), with statistical significance determined through distributional assumption testing, adaptive omnibus tests, and Bonferroni-adjusted pairwise comparisons. For the three static datasets, single train–test splits suggested moderate LSTM performance (C-index: 0.65–0.70); however, repeated evaluation showed that this finding was not robust. Across 100 iterations, CoxPH and RSF consistently outperformed LSTM (all p<0.001), achieving mean C-index values ranging from 0.66 to 0.73 compared with 0.31 to 0.42 for LSTM, with very large effect sizes (Cohen’s d: 6–30). In contrast, on the longitudinal PBC2 dataset, the LSTM-based model achieved the highest repeated C-index (0.806±0.040), compared with 0.777±0.042 for CoxPH and 0.779±0.040 for RSF, with only 1.2% performance degradation under repeated evaluation. These findings indicate that reliance on a single train–test split can produce unstable and potentially unrepresentative estimates of model performance. Traditional survival models were more accurate and stable for datasets containing only static baseline covariates, whereas the LSTM-based model showed higher discrimination on longitudinal survival data with genuine temporal structure, though this advantage is confounded with greater access to patient history and cannot be attributed to architecture alone. Overall, the results underscored the importance of repeated evaluation as a more reliable framework for comparing survival models, and are consistent with, though do not conclusively establish, the importance of matching model architecture to data characteristics.