Machine Learning-Based Prediction of Construction Project Costs and Overruns Using a Comparative Study of Benchmark Datasets
Abstract
A common failure in construction project delivery around the world is cost overruns and schedule delays. This research benchmarks five machine learning models (Random Forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Support Vector Machine (SVM), and a Multilayer Perceptron Artificial Neural Network (ANN)) on two public benchmark datasets, under a single unified protocol: (1) the UCI Residential Building Data Set consisting of 107 predictors and 372 projects for construction cost and sales price prediction, and (2) 2,590 project-phase records from the New York City (NYC) Open Data portal, for time and cost overrun prediction and delay-risk classification. All models were tuned by randomized search with 5-fold cross-validation and evaluated on held-out test sets, and the best models were interpreted using SHapley Additive exPlanations (SHAP). XGBoost achieved the best cost-prediction performance on the UCI benchmark (R² = 0.978, MAPE = 8.74%), exceeding recently published results, with SHAP attributing predictions primarily to the preliminary cost estimate, construction duration, and lagged macro-economic indices. On the NYC data, point regression of overrun magnitude from administrative features alone proved weak (best R² = 0.32), whereas binary delay-risk classification was strong (LightGBM F1 = 0.734, AUC = 0.829) on a nearly balanced sample, matching or exceeding prior questionnaire-based delay classifiers. The findings indicate that gradient-boosted ensembles are the most robust default for heterogeneous construction data, and that machine learning can already serve as an effective delay-triage tool using routinely archived administrative records — a practice directly transferable to public project governance in developing contexts such as Iraq.