Skip to content
Review Open access

Explainable Machine Learning for Compressive Stress Prediction in Cement-Based Construction Materials: A Systematic Benchmarking and SHAP Interpretability Study

Sep 2026 · Electronic Journal of Structural Engineering · 0 citations

Abstract

Accurate prediction of the compressive strength of cement-based materials remains a central problem in structural engineering, because the property depends on strongly nonlinear interactions among mixture proportions, supplementary cementitious materials and curing age. Recent systematic reviews of the machine learning literature in this area report three recurring weaknesses: model comparisons are seldom made under one shared data split and are rarely validated statistically; raw mix-design variables are usually passed to the learner without encoding established concrete science; and analysis code is almost never released, so published accuracies cannot be reproduced or placed on a common footing. This study addresses all three on the University of California, Irvine (UCI) Concrete Compressive Strength dataset (n = 1,030). Nine regression models, from ordinary least squares to gradient-boosted ensembles, were compared under one identical ten-fold cross-validation (CV) protocol on a fixed 80/20 stratified split, with all model selection confined to the training partition. The eight raw mix-design variables were augmented with eight physics-informed features derived from Abrams’ strength law and cement hydration kinetics, among them the water-to-cement ratio, the water-to-binder ratio and log-transformed curing age. Hyperparameters of the two gradient-boosted learners were tuned by Bayesian optimisation in Optuna over 100 trials each. Extreme gradient boosting (XGBoost) gave the best held-out result, with a coefficient of determination (R²) of 0.9567, a root mean squared error (RMSE) of 3.46 MPa and a mean absolute error (MAE) of 2.35 MPa. The four gradient-boosted variants, tuned and untuned XGBoost and light gradient boosting machine (LightGBM), could not be separated from one another by Wilcoxon signed-rank tests, and all four were significantly better than the kernel, bagging and linear baselines (p = 0.002). Bayesian tuning raised the cross-validated score of both boosters yet lowered their held-out score; this negative result is reported in full because it carries a concrete rule for practitioners. SHapley Additive exPlanations (SHAP) computed for both tuned boosters identified curing age, the water-to-binder ratio and total binder content as the dominant predictors, in agreement with physical theory, and the two models returned the same six leading features with closely matching global importance (Spearman ρ = 0.829). The complete pipeline is released as open source under a single fixed random seed.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.