Skip to content
Open access

A Robust Heterogeneous Ensemble Framework for Software Cost Estimation with Prediction Uncertainty Quantification

Aug 2026 · UHD Journal of Science and Technology · 0 citations · 28 references

TL;DR

This paper presents a method using an uncertainty-aware heterogeneous ensemble of 100 bootstrap-trained base learners, which jointly produce point predictions and prediction intervals using a robust trimmed mean, with effort modeled on the log scale.

Abstract

Successfully estimating software costs is critical to project management because poor estimates lead to overruns, scheduling issues, and project failure. Machine learning improves estimates, but usually only point estimates are produced without providing valid uncertainty around them. This paper presents a method using an uncertainty-aware heterogeneous ensemble of 100 bootstrap-trained base learners (using gradient boosting, extra trees, and random forests), which jointly produce point predictions and prediction intervals using a robust trimmed mean, with effort modeled on the log scale. Using the NASA93 dataset (with 93 projects, split 80/20 into 74 for the training set and 19 for the test set) and repeated cross-validation (feature selection and scaling were performed only within the training folds to avoid leakage), the model achieved a mean absolute error of 309.1 and a percentage of relative error deviation from the predicted value rate of 51.3% (30 out of 74 projects) outperforming all eight of the Bayesian baselines it was compared to, while being the only model to provide the validity of its prediction intervals through calibration. Given that normality of the ensemble predictions is rejected, empirical and distribution-free percentiles (with 89.5% of the prediction intervals containing the test project true outcomes) are used because they are not biased by the training data. The average interval width covered by the validation of those prediction intervals to cross-validate the interval width over 86.1% of the test set differs significantly from the linear bases using a paired Wilcoxon test (P < 0.001). Thus, the framework described couples competitive levels of accuracy with quantifiable and empirically validated confidence levels.

Read PDF

Similar papers

Review

Evaluation of Software Effort Estimation Methods for Machine learning techniques

A review of deferent machine learning methods that are using for effort estimation like regression models, decision trees, random forest, neural networks, and then evaluate this models based on performance criteria such as MAE (Mean Absolute Error) and R 2 Score.

Montaser Fadulalla Ahmed Adam, Haroun Abdalla Eissa · 0 citations
#small language model Open access Sep 2026

Comparative Benchmark of Eleven Regression Models for Software Effort Estimation on a COCOMO-Like Dataset

This study develops a comparative benchmark for software effort estimation using the benchmark suite implemented in Python and the result package generated by that suite. Eleven regressors were compared under a leakage-safe protocol on a COCOMO-like dataset of 62 projects and 20 numeric predictors, with 49 projects res...

Jaime Aguilar-Ortiz, Víctor M. Zamudio-García, Marcos Yamir Gómez-Ramos et al. · 0 citations
Open access Aug 2026

Improving Software Effort Estimation Through Feature Selection and Optimized SVR

Accurate software development effort estimation is essential but often hindered by high-dimensional data and the inefficiencies of handling feature selection and parameter tuning as separate, sequential processes. This study proposes an integrated Whale Optimization Algorithm–Support Vector Regression (WOA-SVR) framewo...

R. Putri, G. E. Yuliastuti, Citra Nurina Prabiantissa · 0 citations
Open access Sep 2026

Machine Learning-Based Prediction of Construction Project Costs and Overruns Using a Comparative Study of Benchmark Datasets

A common failure in construction project delivery around the world is cost overruns and schedule delays. This research benchmarks five machine learning models (Random Forest (RF), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Support Vector Machine (SVM), and a Multilayer Perceptron A...

Mustafa Al-Saadi · 0 citations
Open access Sep 2026

Comparative Analysis of Ensemble Learning Methods for Software Reliability Prediction

The results demonstrate that ensemble methods provide superior performance in identifying reliability levels, and which classification method is most suitable for predicting software reliability based on code metrics such as Cyclomatic Complexity and Halstead Volume.

Nadir Subaşı, Ö. Özer · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.