Probabilistic and Interpretable Machine Learning Framework for Predicting Pile Unit Base Resistance in Soft Soil
Accurate prediction of pile base resistance is essential for the safe and economical design of deep foundations, particularly in soft soils where load-transfer mechanisms are highly nonlinear and uncertain. This study develops a comparative, probabilistic, and interpretable machine learning framework for predicting pile unit base resistance using five input variables: applied load, settlement, effective pile length, axial stiffness, and SPT value. A Gaussian Process Regression model with an automatic relevance determination (ARD) Exponential kernel achieved the best performance, with RMSE = 262.11 kPa, R2 = 0.943 on an independent test set, and 95% prediction intervals with 96.46% coverage. Beyond record-level evaluation, a leave-one-pile-out validation (the first grouped validation applied to this database) showed harder generalization to entirely unseen piles, driven mainly by a per-pile level offset rather than shape mismatch (within-pile correlation = 0.975). A sequential next-stage scheme, calibrating this level from a pile’s early loading stages, then predicted its remaining segments with consistently strong agreement (Willmott’s d = 0.76–0.83), supporting practical extension of partial load tests. Interpretability was assessed using ARD, SHAP, permutation/ablation importance, and partial dependence/accumulated local effects analysis, identifying settlement as the dominant predictor. The framework combines accuracy, calibrated uncertainty, interpretability, and validated segment-level extrapolation for reliability-oriented pile assessment.