Skip to content
Open access

Benchmarking Machine Learning Models for Software Project Schedule Overrun Prediction

Jul 2026 · Journal of Computer Science and Technology Studies · Vol 8, pp. 159-212 · 0 citations

TL;DR

Considering both prediction error and interpretability, Gradient Boosting Machine provided the strongest overall balance and represents a practical candidate for early-warning systems, risk-based project prioritization, schedule-recovery planning, and portfolio-level decision support.

Abstract

Software projects frequently exceed their planned schedules because project performance is shaped by interacting technical, organizational, resource, risk, and stakeholder-related factors. This study benchmarks six regression algorithms—Linear Regression, Regression Tree, Random Forest, Support Vector Regression, Gradient Boosting Machine, and Deep Neural Network—to predict the percentage of schedule overrun in software projects and identify the model that provides the most favorable balance between predictive accuracy and interpretability. All six models were developed and evaluated using a publicly available dataset containing 4,517 software-project records. To ensure a consistent comparison, the models used identical predictors, default hyperparameter configurations, and a 10-fold cross-validation procedure. Predictive performance was assessed using root mean squared error, mean absolute error, correlation, and squared correlation. Gradient Boosting Machine achieved the lowest prediction errors, with an RMSE of 4.578 and an MAE of 3.503. Deep Neural Network produced nearly equivalent performance and recorded the highest squared correlation of 0.875. Random Forest ranked third, whereas Support Vector Regression produced the weakest performance under the adopted default settings. Interpretability analysis further identified Risk Factor, Duration Months, and Client Satisfaction Rating as the most consistently influential predictors across the interpretable models. The findings demonstrate the potential of ensemble and deep-learning methods to estimate the magnitude of schedule overrun using structured software-project data. Considering both prediction error and interpretability, Gradient Boosting Machine provided the strongest overall balance and represents a practical candidate for early-warning systems, risk-based project prioritization, schedule-recovery planning, and portfolio-level decision support. Future research should validate the models using independent organizational datasets, optimize their hyperparameters, and apply consistent explainability methods across all algorithms.

Read PDF

Similar papers

#small language model Open access Sep 2026

Comparative Benchmark of Eleven Regression Models for Software Effort Estimation on a COCOMO-Like Dataset

This study develops a comparative benchmark for software effort estimation using the benchmark suite implemented in Python and the result package generated by that suite. Eleven regressors were compared under a leakage-safe protocol on a COCOMO-like dataset of 62 projects and 20 numeric predictors, with 49 projects res...

Jaime Aguilar-Ortiz, Víctor M. Zamudio-García, Marcos Yamir Gómez-Ramos et al. · 0 citations
Open access Sep 2026

Machine Learning-Based Prediction of Construction Project Costs and Overruns Using a Comparative Study of Benchmark Datasets

A common failure in construction project delivery around the world is cost overruns and schedule delays. This research benchmarks five machine learning models (Random Forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Support Vector Machine (SVM), and a Multilayer Perceptron A...

Mustafa Al-Saadi · 0 citations
Open access Aug 2026

A Robust Heterogeneous Ensemble Framework for Software Cost Estimation with Prediction Uncertainty Quantification

This paper presents a method using an uncertainty-aware heterogeneous ensemble of 100 bootstrap-trained base learners, which jointly produce point predictions and prediction intervals using a robust trimmed mean, with effort modeled on the log scale.

H. Sharif, Tara Nawzad Ahmad Al Attar, D. Rashid · 0 citations
Review

Evaluation of Software Effort Estimation Methods for Machine learning techniques

A review of deferent machine learning methods that are using for effort estimation like regression models, decision trees, random forest, neural networks, and then evaluate this models based on performance criteria such as MAE (Mean Absolute Error) and R 2 Score.

Montaser Fadulalla Ahmed Adam, Haroun Abdalla Eissa · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.