Skip to content
Open access

Prediction And Analysis Of Concrete Compressive Strength By Machine Learning Methods

Jun 2026 · Bitlis Eren Üniversitesi Fen Bilimleri Dergisi · Vol 15, pp. 669-685 · 0 citations · 28 references

TL;DR

This work presents a systematic multi-model comparison within a unified hyperparameter optimization framework and highlights that boosting-based machine learning models not only achieve high accuracy but also provide interpretable and robust predictions when evaluated through comprehensive error and explainability analyses.

Abstract

Concrete compressive strength (CCS) is a critical parameter directly affecting the load-bearing capacity, durability, and overall safety of engineering structures. Traditional experimental approaches for determining CCS are time-consuming and costly, making predictive models an attractive alternative. In this study, thirteen different machine learning algorithms were applied to a well-established dataset (1030 samples, 8 input parameters) to estimate concrete compressive strength. Unlike many previous studies using the Yeh dataset that primarily emphasize prediction accuracy of individual models, this work presents a systematic multi-model comparison within a unified hyperparameter optimization framework. In addition to conventional performance metrics, permutation importance and SHAP-based explainability analyses are jointly employed, and detailed error evaluations are conducted across curing age and water-to-binder ratio subgroups to enhance engineering interpretability. Among the models tested, the CatBoost algorithm demonstrated the highest predictive performance (R² = 0.9469, RMSE = 3.70), followed closely by XGBoost, Gradient Boosting, and a stacking ensemble model. The results highlight that boosting-based machine learning models not only achieve high accuracy but also provide interpretable and robust predictions when evaluated through comprehensive error and explainability analyses.

Read PDF

Similar papers

Open access Jul 2026

Machine Learning Prediction of Concrete Compressive Strength: Model Comparison, CatBoost Optimization, and SHAP Interpretation

A comparative framework evaluating nine regression algorithms using the UCI Concrete Compressive Strength dataset, jointly integrating correlation-corrected statistical validation, multi-model Bayesian optimization, and domain-informed feature engineering with SHAP interpretation, rarely combined in prior concrete-strength studies.

Musthafa 'Abduh Fakhruddin, Sri Winarno, Acun Kardianawati · 0 citations
Conference Open access Jul 2026

Machine Learning-Based Prediction of Compressive Strength in Basalt Fiber Reinforced Concrete

Accurate prediction of the mechanical strength of Basalt Fiber Reinforced Concrete (BFRC) is critical for structural design, safety assessment, and the advancement of sustainable infrastructure in civil engineering. Traditional prediction methods often fail to capture the nonlinear relationships between BFRC mix proportions and resulting strength characteristics, leading to unreliable estimations. To address this limitation, this study proposes the Optimized Moment Balanced Machine (OMBM), an advanced machine learning model developed to improve the predictive accuracy of BFRC strength parameters. The model was trained and evaluated using key input features, including cement content, silica fume, fly ash, superplasticizer, water, aggregate composition, and fiber property parameters. The performance of the OMBM was benchmarked against four established machine learning models, such as Least Squares Support Vector Machine (LSSVM), Backpropagation Neural Network (BPNN), K-Nearest Neighbors (KNN), and Linear Regression (LR). Results from ten-fold cross-validation show that OMBM consistently outperforms the comparison models across five evaluation metrics. It achieved the lowest RMSE (2.411), MAE (1.788), and MAPE (4.08%), along with the highest values for correlation coefficient (R = 0.978), and coefficient of determination (R2 = 0.956). Furthermore, the OMBM achieved a Reference Index (RI) score of 1.000, which confirms its position as the leading predictive model within this comparative framework. These results confirm the robustness and reliability of the proposed OMBM model, making it a highly effective tool for accurate strength prediction of BFRC. This approach offers significant potential for the advancement of sustainable infrastructure by enabling more accurate and efficient use of concrete materials.

R. R. Khasani, Ferry Hermawan, Yuliana Usman · 0 citations
Open access Jul 2026

Machine Learning-Based Compressive Strength Prediction and Multi-Objective Optimization of Ultra-High Performance Concrete

The compressive strength of ultra-high-performance concrete (UHPC) is jointly influenced by multiple factors, including material composition, mixture proportion parameters, and curing regime. Conventional empirical methods are therefore insufficient to accurately characterize the highly nonlinear relationships involved. To improve the prediction accuracy of UHPC compressive strength and to achieve mixture proportion optimization that simultaneously considers mechanical performance, economic efficiency, and environmental impact, this study developed random forest (RF), artificial neural network (ANN), gradient boosting decision tree (GBDT), and extreme gradient boosting (XGBoost) models based on 810 publicly available UHPC experimental datasets. Model performance was evaluated using R2, RMSE, MAE, and MAPE. To enhance the robustness of model validation, repeated K-fold cross-validation, sensitivity analysis with different random seed splits, and benchmark model comparisons were further introduced. The results indicate that the XGBoost model achieved superior predictive performance on both the test set and robustness validation, with test-set R2, RMSE, MAE, and MAPE values of 0.9604, 7.77, 5.58, and 4.80, respectively. The model was further interpreted using SHAP, PDP, and ICE methods, and the results revealed that curing age, fiber content, silica fume content, and water-to-binder ratio were important variables affecting the compressive strength of UHPC. Furthermore, XGBoost was used as a surrogate model and coupled with NSGA-II and TOPSIS methods for multi-objective optimization. Under the constraints of compressive strength, water-to-binder ratio, superplasticizer-to-binder ratio, and absolute volume, a computationally recommended UHPC mixture proportion balancing strength, cost, and carbon emissions was obtained. This study provides a reproducible machine-learning-assisted approach for UHPC compressive strength prediction and low-carbon, cost-effective mixture proportion design.

Rong Li, Teng Zhou, Siyu Lu et al. · 0 citations
Open access Aug 2026

Comparative machine learning models for predicting the compressive strength of ultra-high-performance concrete

Ultra-high-performance concrete (UHPC) exhibits exceptional mechanical properties and durability. However, its compressive strength is highly dependent on complex mix design parameters. While traditional experimental techniques and regression-based models are commonly used to evaluate UHPC compressive strength, machine learning approaches offer an efficient alternative for capturing complex nonlinear relationships. This study develops a machine learning–based framework to predict the compressive strength of UHPC and compares the predictive performance of five advanced algorithms: Extremely Randomized Trees (ER), Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), CatBoost, and Artificial Neural Network (ANN). A comprehensive experimental database was utilized for training and validation purposes. Among the evaluated models, CatBoost achieved the best predictive performance, with a coefficient of determination (R²) exceeding 0.90, a root mean square error (RMSE) of approximately 4.5 MPa, and a mean absolute error (MAE) of approximately 3.6 MPa. However, subgroup residual analysis showed that the prediction reliability was not uniform across the full strength range. In particular, mixtures with compressive strength ≥180 MPa exhibited larger errors and systematic underprediction, mainly due to the limited number of ultra-high-strength samples in the compiled database. Therefore, the model is more reliable within well-represented strength ranges, while predictions in the ultra-high-strength region should be interpreted with caution. SHAP-based analysis, feature dependency analysis, and both Individual Conditional Expectation (ICE) and Partial Dependence Plots (PDP) were employed. These explainable AI techniques identified key variables and quantified their contributions to the compressive strength of UHPC. The findings demonstrate that interpretable machine learning can support preliminary UHPC mixture assessment by combining predictive performance with physically meaningful insights.

Nga T. T. Nguyen, Tuan Anh Nguyen, Tu Tuan Nguyen et al. · 0 citations
Open access Aug 2026

Dominant-learner adaptive mixing for concrete compressive strength prediction

Accurate prediction of concrete compressive strength is essential for mixture design, quality control, and the broader use of supplementary cementitious materials in low-carbon construction. Fly ash concrete is particularly challenging to model because its strength development is affected by nonlinear interactions among binder composition, water–binder relationships, admixture dosage, and material characteristics. To address this problem, this study proposes a Dominant Learner with Adaptive Mixing (DLAM) framework for data-driven strength prediction. DLAM uses inner cross-validation to identify the most reliable learner from a pool of machine learning models and introduces a validation-controlled Ridge calibration step to exploit complementary information among candidate predictions. The calibration branch is adopted only when it improves the inner-validation root mean squared error (RMSE), thereby reducing the risk of unnecessary model combination and performance degradation. The framework is evaluated using a leakage-free repeated outer/inner validation protocol on a fly ash concrete dataset and is further examined on an independent public concrete strength dataset. DLAM is compared with individual learners, adaptive model-averaging baselines, and Stacking. The results show that DLAM achieves the lowest mean RMSE among the focused comparators on both datasets, with a clear improvement on the external dataset and a more modest gain on the fly ash dataset. These findings demonstrate that validation-controlled calibration provides a transparent and robust way to enhance machine-learning-based concrete strength prediction, especially when different learners capture complementary aspects of the mixture–strength relationship.

Jinjin Wang, Zhihao Zhao, Mingjie Han · 0 citations
Open access Aug 2026

Novel Interpretable Machine Learning Models for Predicting Compressive Strength of Nano-Silica Concrete

This study presents a comprehensive comparative analysis of several machine learning (ML) models for predicting the compressive strength (CS) of nano-silica (NS)-enhanced concrete. A large dataset comprising 724 experimental mix designs was compiled from the literature, including various variables such as cement content, water-to-binder ratio, fine and coarse aggregates, nano-silica content, superplasticizer content, and curing time. Six ML algorithms were developed and evaluated: Interaction Model, Full Quadratic (FQ), Artificial Neural Network (ANN), M5P-Tree, Gradient Boosting (GB), and Random Forest (RF). Model performance was assessed using R², RMSE, MAE, scatter index (SI), and objective value (OBJ). Among all models, the RF model achieved the highest predictive accuracy, followed by ANN and GB models. Sensitivity analysis revealed curing time as the most influential factor, while partial dependence plots exhibited the nonlinear effect of nano-silica quantity, with an optimal strength response about 15 kg/m3. In addition, SHAP (SHapley Additive exPlanations) analysis was employed to enhance model interpretability, confirming the dominant influence of curing age and water-to-cement ratio on compressive strength prediction. Compared with many previous studies relying on limited datasets or single-model approaches, this study provides a robust, interpretable, and generalizable ML framework for optimizing nano-silica concrete mix design. The findings highlight the strong potential of ML, particularly ensemble models combined with explainable AI techniques, to improve prediction reliability, reduce trial-and-error experimentation, and support more cost-efficient and sustainable concrete design.

Yousif J. Bas, Jamal I. Kakrasul, Kamaran S. Ismail et al. · 0 citations