Skip to content
Open access

Comparative Analysis of AI and Statistical Models for Predicting Mechanical and Durability-Related Properties of Alkali-Activated Recycled Aggregate Concrete

Jul 2026 · Buildings · 0 citations · 46 references

Abstract

Alkali-activated recycled aggregate concrete (AARAC) offers a sustainable alternative to traditional concrete but suffers from complex, non-linear mechanical behavior that challenges conventional prediction methods. This study develops and compares five machine learning models, linear regression (LR), M5P, Random Forest (RF), K-Nearest Neighbors (KNN) and XGBoost, for predicting the compressive strength (Cs), flexural strength (Fs), splitting tensile strength (Ss), pull-out bond strength (PT), and water absorption (Wa%) of AARAC. A dataset of 360 experimental samples, incorporating natural aggregate, recycled concrete aggregate (RCA), cement block aggregate (CBA), water-to-cement ratio (W/C), alkaline treatment status, and slump, was used. Models were evaluated via train/test split (80/20) and 10-fold cross-validation using R2, MAE, RMSE, and MAPE. Random Forest achieved the highest test R2 (0.8736) and lowest test MAPE (1.418%) and XGBoost (R2 = 0.8605, MAPE = 1.557%). KNN and M5P performed moderately, while LR was the weakest (R2 = 0.6958, MAPE = 2.147%). All tree-based models exhibited overfitting, with training R2 up to 0.98. Scatter plot analysis revealed systematic underprediction by RF for Cs (constant offset of ~2 MPa) and increasing bias for PT, Ss, and Wa% at higher values. XGBoost gave perfect predictions for PT and Wa% but underpredicted Cs and Fs. K-fold cross-validation confirmed XGBoost as the most robust (mean R2 = 0.9844). Correlation analysis showed W/C strongly increases Wa% (r = 0.80) and decreases PT (r = −0.73); RCA negatively affects mechanical properties, while CBA and alkaline treatment improve them. The study concludes that ensemble tree models, particularly Random Forest, are superior for AARAC prediction, but systematic bias requires post hoc calibration.

Read PDF

Similar papers

Open access Jul 2026

Data-driven modeling of compressive strength in sustainable self-compacting concrete incorporating recycled aggregates using ensemble learning techniques

Abstract This study develops a robust framework for estimating the compressive strength of self-compacting concrete (SCC) incorporating recycled aggregates using supervised machine learning (ML) techniques. A comprehensive experimental database comprising 582 concrete mix designs was used, encompassing diverse input variables including binder content, water, coarse and fine aggregates, recycled aggregate proportion, superplasticizer dosage, and curing time. Seven ML algorithms—XGBoost, CatBoost, AdaBoost, Extra Trees, Bagging Regressor, K-Nearest Neighbors, and Radius Neighbors—were systematically trained using a stratified 70/15/15 data split and optimized via grid search with five-fold cross-validation. Model performance was evaluated using coefficient of determination (R 2), root mean squared error, and MAE across training, validation, and testing datasets. Among all models, XGBoost demonstrated the highest accuracy, achieving an average R 2 of 0.9799, RMSE of 2.87 MPa, and mean absolute error of 1.97 MPa. The Permutation Feature Importance analysis revealed that binder content, water, and coarse aggregate were the most influential predictors of strength. This study confirms that ensemble ML models, particularly XGBoost, can reliably predict the compressive strength of SCC with recycled aggregates, while offering transparent insights into material behavior. The results provide a valuable tool for sustainable mix design optimization and practical implementation in eco-efficient concrete construction.

A. Khan, M. D. Rasheed, Muhammad Huzaifa Naveed et al. · 0 citations
Open access Aug 2026

Novel Interpretable Machine Learning Models for Predicting Compressive Strength of Nano-Silica Concrete

This study presents a comprehensive comparative analysis of several machine learning (ML) models for predicting the compressive strength (CS) of nano-silica (NS)-enhanced concrete. A large dataset comprising 724 experimental mix designs was compiled from the literature, including various variables such as cement content, water-to-binder ratio, fine and coarse aggregates, nano-silica content, superplasticizer content, and curing time. Six ML algorithms were developed and evaluated: Interaction Model, Full Quadratic (FQ), Artificial Neural Network (ANN), M5P-Tree, Gradient Boosting (GB), and Random Forest (RF). Model performance was assessed using R², RMSE, MAE, scatter index (SI), and objective value (OBJ). Among all models, the RF model achieved the highest predictive accuracy, followed by ANN and GB models. Sensitivity analysis revealed curing time as the most influential factor, while partial dependence plots exhibited the nonlinear effect of nano-silica quantity, with an optimal strength response about 15 kg/m3. In addition, SHAP (SHapley Additive exPlanations) analysis was employed to enhance model interpretability, confirming the dominant influence of curing age and water-to-cement ratio on compressive strength prediction. Compared with many previous studies relying on limited datasets or single-model approaches, this study provides a robust, interpretable, and generalizable ML framework for optimizing nano-silica concrete mix design. The findings highlight the strong potential of ML, particularly ensemble models combined with explainable AI techniques, to improve prediction reliability, reduce trial-and-error experimentation, and support more cost-efficient and sustainable concrete design.

Yousif J. Bas, Jamal I. Kakrasul, Kamaran S. Ismail et al. · 0 citations
Open access Aug 2026

Explainable Random Forest Framework for Predicting Compressive Strength of Sustainable Concrete Incorporating Industrial Waste Materials

Compressive strength is the single most important design parameter governing the safety, serviceability, and economy of concrete structures, yet its determination through standard 7-, 14-, or 28-day destructive cylinder/cube testing is slow, costly, and unable to assess concrete already cast in place. This study develops and evaluates a Random Forest (RF) regression model to predict the compressive strength of concrete directly from eight standard mix-design parameters — cement, blast furnace slag, fly ash, water, superplasticizer, coarse aggregate, fine aggregate, and curing age — using Yeh's (1998) benchmark dataset of 1,030 experimentally tested concrete mixtures. Following data cleaning, exploratory correlation analysis, an 80:20 train-test split, and five-fold GridSearchCV hyperparameter tuning, the optimized Random Forest model is benchmarked against Linear Regression, Ridge Regression, and Support Vector Regression using the coefficient of determination (R²), Root Mean Squared Error (RMSE), and Mean Absolute Error (MAE). The Random Forest model achieves the strongest predictive performance of the models tested, substantially outperforming the linear baselines and confirming that concrete strength development is governed by non-linear interactions among mix constituents. Feature importance analysis further shows that curing age and cement content are the dominant predictors, while water content exerts a clear negative influence consistent with Abrams' Law, and coarse/fine aggregates contribute comparatively little, consistent with their role as largely inert fillers. These findings demonstrate that Random Forest regression offers a fast, accurate, and interpretable, non-destructive alternative to conventional strength testing, with practical value for mix-design optimization, quality control, and early-stage structural decision-making.

M. Selvakumar, S. Geetha, P. Krishna Kumar et al. · 0 citations
Open access Jul 2026

Prediction of Fresh and Mechanical Properties of Self-Compacting Concrete Using Gaussian Process Regression

Self-compacting concrete (SCC) incorporating ceramic waste powder (CWP) as a partial cement replacement offers a route to reduce both landfill burden and embodied carbon in construction, but characterising the combined effect of mix design variables on the full suite of fresh and hardened properties normally demands extensive laboratory testing. This study evaluates Gaussian Process Regression (GPR) as a data-efficient surrogate modelling approach for predicting twelve fresh and mechanical properties of CWP-based SCC from three mix design inputs: cement content, CWP content, and water-to-binder (w/b) ratio. A dataset of twenty-one experimentally characterised SCC mixes, spanning slump flow, T500 time, V-funnel time, L-box ratio, segregation resistance, compressive strength (7, 28, and 90 days), split tensile strength (7, 28, and 90 days), and 28-day flexural strength, was modelled using independent GPR models with a Matern 5/2 kernel and white-noise term, with performance assessed exclusively through leave-one-out cross-validation (LOOCV) given the small sample size. Eleven of the twelve properties were predicted with LOOCV coefficients of determination (R2 ) between 0.80 and 0.99, with the strongest performance observed for slump flow and split tensile strength (R2 ≥ 0.96). Twenty-eight-day flexural strength was the exception, with a negative LOOCV R2 indicating that the three mix design inputs alone do not explain its variability in this dataset. The results demonstrate that GPR, combined with rigorous LOOCV validation and explicit predictive uncertainty, is a viable and transparent tool for mix-design-stage property prediction in small experimental SCC datasets, while also illustrating the diagnostic value of LOOCV in identifying properties that require additional explanatory variables.

P. Joshi, D. Parekh, Parth Harkishan et al. · 0 citations
Conference Jul 2026

Machine Learning–Based Prediction of Mechanical Properties of Sustainable Concrete

There has been a rise in the need of concrete leading to high consumption of cement and emission of carbon. Sustainable concrete incorporating supplementary cementitious materials (SCMs) is an alternative that is eco-friendly, but the mechanical behavior is complicated and hard to predict using the traditional tests. The work presents a framework involving machine learning because of predicting the mechanical properties of sustainable concrete, namely, compressive, split tensile, and flexural strength. Four models such as Linear Regression, Support Vector Regression, Random Forest, and Artificial Neural Network were constructed based on a data of nearly 1000 sustainable concrete mixes. To estimate the model performance, R 2, RMSE and MAE were used. The findings indicated that ANN and RF had the greatest prediction accuracy. The feature analysis established the most influential factors to be water content, cement dosage, replacement ratio of SCM, and curing age. The mix design approach proposed here is a fast, economical, and sustainable approach to the design of concrete mix.

Arti Chouksey, Santosh Reddy P, P. S et al. · 0 citations
Conference Open access Jul 2026

Machine Learning-Based Prediction of Compressive Strength in Basalt Fiber Reinforced Concrete

Accurate prediction of the mechanical strength of Basalt Fiber Reinforced Concrete (BFRC) is critical for structural design, safety assessment, and the advancement of sustainable infrastructure in civil engineering. Traditional prediction methods often fail to capture the nonlinear relationships between BFRC mix proportions and resulting strength characteristics, leading to unreliable estimations. To address this limitation, this study proposes the Optimized Moment Balanced Machine (OMBM), an advanced machine learning model developed to improve the predictive accuracy of BFRC strength parameters. The model was trained and evaluated using key input features, including cement content, silica fume, fly ash, superplasticizer, water, aggregate composition, and fiber property parameters. The performance of the OMBM was benchmarked against four established machine learning models, such as Least Squares Support Vector Machine (LSSVM), Backpropagation Neural Network (BPNN), K-Nearest Neighbors (KNN), and Linear Regression (LR). Results from ten-fold cross-validation show that OMBM consistently outperforms the comparison models across five evaluation metrics. It achieved the lowest RMSE (2.411), MAE (1.788), and MAPE (4.08%), along with the highest values for correlation coefficient (R = 0.978), and coefficient of determination (R2 = 0.956). Furthermore, the OMBM achieved a Reference Index (RI) score of 1.000, which confirms its position as the leading predictive model within this comparative framework. These results confirm the robustness and reliability of the proposed OMBM model, making it a highly effective tool for accurate strength prediction of BFRC. This approach offers significant potential for the advancement of sustainable infrastructure by enabling more accurate and efficient use of concrete materials.

R. R. Khasani, Ferry Hermawan, Yuliana Usman · 0 citations