Jul 2026· IDEALIS : InDonEsiA journaL Information System· 0 citations· 45 references
TL;DR
A comparative framework evaluating nine regression algorithms using the UCI Concrete Compressive Strength dataset, jointly integrating correlation-corrected statistical validation, multi-model Bayesian optimization, and domain-informed feature engineering with SHAP interpretation, rarely combined in prior concrete-strength studies.
Abstract
Accurate prediction of concrete compressive strength is vital for structural design, yet conventional testing is constrained by lengthy curing requirements. Machine learning offers an alternative by modeling non-linear mix-performance interactions. This study presents a comparative framework evaluating nine regression algorithms using the UCI Concrete Compressive Strength dataset (1,005 samples). Performance was assessed via 10x5 repeated cross-validation with 95% confidence intervals, and statistical significance was evaluated using a Linear Mixed-Effects Model with Holm-Bonferroni corrected pairwise t-tests. Tree-based ensembles outperformed linear approaches, with CatBoost yielding the highest baseline cross-validation R² of 0.931 (95% CI: 0.927 to 0.935). Subsequent Bayesian hyperparameter optimization via Optuna’s Tree-structured Parzen Estimator (400 trials) improved the final CatBoost model’s performance to a test of 0.943, RMSE of 4.142 MPa, and MAE of 2.616 MPa. SHAP analysis indicated that curing age, the water-to-binder ratio, and cement are the dominant predictors, while the model exhibited physically consistent behavior aligned with concrete hydration kinetics and Abrams' law. This work's key contribution is jointly integrating correlation-corrected statistical validation, multi-model Bayesian optimization, and domain-informed feature engineering with SHAP interpretation, rarely combined in prior concrete-strength studies. The framework offers an accurate, interpretable tool for preliminary concrete mix design.
The compressive strength of ultra-high-performance concrete (UHPC) is jointly influenced by multiple factors, including material composition, mixture proportion parameters, and curing regime. Conventional empirical methods are therefore insufficient to accurately characterize the highly nonlinear relationships involved. To improve the prediction accuracy of UHPC compressive strength and to achieve mixture proportion optimization that simultaneously considers mechanical performance, economic efficiency, and environmental impact, this study developed random forest (RF), artificial neural network (ANN), gradient boosting decision tree (GBDT), and extreme gradient boosting (XGBoost) models based on 810 publicly available UHPC experimental datasets. Model performance was evaluated using R2, RMSE, MAE, and MAPE. To enhance the robustness of model validation, repeated K-fold cross-validation, sensitivity analysis with different random seed splits, and benchmark model comparisons were further introduced. The results indicate that the XGBoost model achieved superior predictive performance on both the test set and robustness validation, with test-set R2, RMSE, MAE, and MAPE values of 0.9604, 7.77, 5.58, and 4.80, respectively. The model was further interpreted using SHAP, PDP, and ICE methods, and the results revealed that curing age, fiber content, silica fume content, and water-to-binder ratio were important variables affecting the compressive strength of UHPC. Furthermore, XGBoost was used as a surrogate model and coupled with NSGA-II and TOPSIS methods for multi-objective optimization. Under the constraints of compressive strength, water-to-binder ratio, superplasticizer-to-binder ratio, and absolute volume, a computationally recommended UHPC mixture proportion balancing strength, cost, and carbon emissions was obtained. This study provides a reproducible machine-learning-assisted approach for UHPC compressive strength prediction and low-carbon, cost-effective mixture proportion design.
Rong Li, Teng Zhou, Siyu Lu et al.· Applied Sciences· 0 citations
The use of supplementary cementitious materials such as fly ash can reduce environmental impacts and improve the sustainability of concrete construction. However, the nonlinear interactions among mixture design parameters make accurate prediction of concrete compressive strength challenging. In this study, TabPFN, a pre-trained foundation model for tabular data, was applied to predict the compressive strength of fly ash concrete and compared with tuned Random Forest, support vector regression, an artificial neural network, LightGBM, CatBoost, Ridge regression, and Abrams empirical regression. A dataset containing 1062 samples and eight mixture-level variables was used for model development and evaluation. Predictive performance was assessed using the coefficient of determination, mean absolute error, and root mean square error over 100 repeated random splits. The results showed that TabPFN achieved the best overall performance, with an average coefficient of determination of 0.9329, a mean absolute error of 3.2758 MPa, and a root mean square error of 4.6678 MPa. Compared with the strongest tuned gradient-boosting baseline, CatBoost, TabPFN reduced the mean absolute error and root mean square error by 0.8768 MPa and 0.8560 MPa, respectively. Furthermore, repeated-split conformal prediction demonstrated reliable uncertainty quantification, with an average prediction interval coverage probability of 0.9615 and a mean prediction interval width of 23.4554 MPa. SHAP analysis identified the water-to-cement ratio, mortar strength, and water-to-binder ratio as important variables, while additional multicollinearity and feature ablation analyses indicated that correlated ratio variables should be interpreted cautiously. The results indicate that TabPFN provides an accurate, robust, and uncertainty-aware framework for preliminary prediction of 28-day fly ash concrete compressive strength.
Zhihao Zhao, Jinjin Wang, Guohui Ma et al.· Buildings· 0 citations
Results suggest that, within the present five-fold cross-validation setting and limited-sample dataset, RBF kernel ridge regression captures the nonlinear relationships more effectively than conventional linear models; however, broader generalization should be verified using larger datasets and additional validation.
Yuchen Lin· International Conference on...· 0 citations
Ultra-high-performance concrete (UHPC) exhibits exceptional mechanical properties and durability. However, its compressive strength is highly dependent on complex mix design parameters. While traditional experimental techniques and regression-based models are commonly used to evaluate UHPC compressive strength, machine learning approaches offer an efficient alternative for capturing complex nonlinear relationships. This study develops a machine learning–based framework to predict the compressive strength of UHPC and compares the predictive performance of five advanced algorithms: Extremely Randomized Trees (ER), Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), CatBoost, and Artificial Neural Network (ANN). A comprehensive experimental database was utilized for training and validation purposes. Among the evaluated models, CatBoost achieved the best predictive performance, with a coefficient of determination (R²) exceeding 0.90, a root mean square error (RMSE) of approximately 4.5 MPa, and a mean absolute error (MAE) of approximately 3.6 MPa. However, subgroup residual analysis showed that the prediction reliability was not uniform across the full strength range. In particular, mixtures with compressive strength ≥180 MPa exhibited larger errors and systematic underprediction, mainly due to the limited number of ultra-high-strength samples in the compiled database. Therefore, the model is more reliable within well-represented strength ranges, while predictions in the ultra-high-strength region should be interpreted with caution. SHAP-based analysis, feature dependency analysis, and both Individual Conditional Expectation (ICE) and Partial Dependence Plots (PDP) were employed. These explainable AI techniques identified key variables and quantified their contributions to the compressive strength of UHPC. The findings demonstrate that interpretable machine learning can support preliminary UHPC mixture assessment by combining predictive performance with physically meaningful insights.
Nga T. T. Nguyen, T. Nguyen, Tuan-Khoi Nguyen et al.· PLoS ONE· 0 citations
Accurate prediction of the mechanical strength of Basalt Fiber Reinforced Concrete (BFRC) is critical for structural design, safety assessment, and the advancement of sustainable infrastructure in civil engineering. Traditional prediction methods often fail to capture the nonlinear relationships between BFRC mix proportions and resulting strength characteristics, leading to unreliable estimations. To address this limitation, this study proposes the Optimized Moment Balanced Machine (OMBM), an advanced machine learning model developed to improve the predictive accuracy of BFRC strength parameters. The model was trained and evaluated using key input features, including cement content, silica fume, fly ash, superplasticizer, water, aggregate composition, and fiber property parameters. The performance of the OMBM was benchmarked against four established machine learning models, such as Least Squares Support Vector Machine (LSSVM), Backpropagation Neural Network (BPNN), K-Nearest Neighbors (KNN), and Linear Regression (LR). Results from ten-fold cross-validation show that OMBM consistently outperforms the comparison models across five evaluation metrics. It achieved the lowest RMSE (2.411), MAE (1.788), and MAPE (4.08%), along with the highest values for correlation coefficient (R = 0.978), and coefficient of determination (R2 = 0.956). Furthermore, the OMBM achieved a Reference Index (RI) score of 1.000, which confirms its position as the leading predictive model within this comparative framework. These results confirm the robustness and reliability of the proposed OMBM model, making it a highly effective tool for accurate strength prediction of BFRC. This approach offers significant potential for the advancement of sustainable infrastructure by enabling more accurate and efficient use of concrete materials.
R. R. Khasani, Ferry Hermawan, Yuliana Usman· IOP Conference Series: Earth...· 0 citations
To address the limitations of traditional BP neural networks in predicting manufactured sand concrete strength, specifically their susceptibility to local optima and “black-box” opacity, this study developed an integrated framework combining improved optimization algorithms with the Shapley Additive Explanations (SHAP) method. Using a dataset of 375 data points, genetic algorithm-back propagation (GA-BP) and GOOSE-BP prediction models were developed, with AutoFeat employed for explicit model construction based on a SHAP feature analysis. The results demonstrate that the GOOSE-BP model significantly outperformed traditional methods, achieving an R2 of 0.916 and reducing prediction errors by 47.5%. The SHAP analysis identified paste thickness and stone powder content as the primary determinants of strength. Key thresholds were established, including a water-to-binder ratio sensitivity range of 0.35–0.50, an optimal stone powder content of 80–110 kg/m3, and a recommended sand ratio of 0.38–0.45. By converting complex nonlinear mappings into interpretable explicit expressions, this study provides a robust scientific basis and a practical computational tool for predicting concrete strength, facilitating the deep integration of machine learning with civil engineering practice.
Juanjuan Quan, Kunlin Liu, Hao Su et al.· Buildings· 1 citation