Skip to content

Comparative Benchmark of Eleven Regression Models for Software Effort Estimation on a COCOMO-Like Dataset

Sep 2026 · International Journal of Combinatorial Optimization Problems and Informatics · Vol 17, pp. 80-105 · 0 citations · 10 references
Software Engineering Research

Abstract

This study develops a comparative benchmark for software effort estimation using the benchmark suite implemented in Python and the result package generated by that suite. Eleven regressors were compared under a leakage-safe protocol on a COCOMO-like dataset of 62 projects and 20 numeric predictors, with 49 projects reserved for model development and 13 for final hold-out testing. The evaluated methods were a three-hidden-layer deep neural network, CatBoost, XGBoost, stacked ensemble regression, random forest, support vector regression with radial basis kernel, Gaussian process regression, LightGBM, an optimized M5P model tree, gradient boosting, and AdaBoost. Model quality was judged through repeated 5×2 cross-validation on the training partition and through an independent hold-out test set, using MAE, RMSE, R², MMRE, MdMRE, and PRED(25). The repeated-CV ranking identified Gaussian process regression as the most accurate and stable model, with mean rank = 1.0000, RMSE = 0.2152, R² = 0.8998, and PRED(25) = 0.9289. The same model also dominated the hold-out evaluation, achieving MAE = 0.1444, RMSE = 0.1835, R² = 0.9084, MMRE = 0.0850, MdMRE = 0.0634, and PRED(25) = 0.9231. The overall order of precision from repeated cross-validation was: Gaussian process regression, optimized model tree (M5P), stacked ensemble regressor, gradient boosting regressor, XGBoost regressor, random forest regressor, CatBoost regressor, AdaBoost regressor, support vector regression, deep neural network, and LightGBM regressor. Permutation importance of the best model showed KDSI, ACAP, PCAP, RELY, and AAF as the most influential predictors. The evidence indicates that, for the present small-sample COCOMO-like setting, kernel-based and piecewise linear/tree-based methods outperform the deeper neural alternative.   Spanish-language metadata / Metadatos en españolTítulo en español:Benchmark comparativo de once modelos de regresión para la estimación del esfuerzo de software en un conjunto de datos similar a COCOMO Resumen:Este estudio desarrolla un benchmark comparativo para la estimación del esfuerzo de software utilizando la batería de evaluación implementada en Python y el paquete de resultados generado por dicha batería. Se compararon once regresores mediante un protocolo diseñado para evitar fugas de información (data leakage) sobre un conjunto de datos similar a COCOMO compuesto por 62 proyectos y 20 predictores numéricos, de los cuales 49 proyectos se reservaron para el desarrollo de los modelos y 13 para la prueba final con un conjunto hold-out. Los métodos evaluados fueron una red neuronal profunda con tres capas ocultas, CatBoost, XGBoost, regresión mediante un ensamble apilado (stacked ensemble), bosque aleatorio, regresión de vectores de soporte con kernel de base radial, regresión mediante procesos gaussianos, LightGBM, un árbol de modelos M5P optimizado, gradient boosting y AdaBoost. La calidad de los modelos se evaluó mediante validación cruzada repetida 5×2 sobre la partición de entrenamiento y mediante un conjunto independiente de prueba hold-out, utilizando MAE, RMSE, R², MMRE, MdMRE y PRED(25). La clasificación obtenida mediante validación cruzada repetida identificó la regresión mediante procesos gaussianos como el modelo más preciso y estable, con un rango medio = 1,0000, RMSE = 0,2152, R² = 0,8998 y PRED(25) = 0,9289. El mismo modelo también dominó la evaluación hold-out, alcanzando MAE = 0,1444, RMSE = 0,1835, R² = 0,9084, MMRE = 0,0850, MdMRE = 0,0634 y PRED(25) = 0,9231. El orden global de precisión obtenido mediante validación cruzada repetida fue el siguiente: regresión mediante procesos gaussianos, árbol de modelos optimizado (M5P), regresor de ensamble apilado, regresor de gradient boosting, regresor XGBoost, regresor de bosque aleatorio, regresor CatBoost, regresor AdaBoost, regresión de vectores de soporte, red neuronal profunda y regresor LightGBM. El análisis de importancia por permutación del mejor modelo mostró que KDSI, ACAP, PCAP, RELY y AAF fueron los predictores más influyentes. La evidencia indica que, para el presente escenario de tamaño muestral reducido y similar a COCOMO, los métodos basados en kernels y los enfoques lineales por tramos o basados en árboles superan a la alternativa neuronal más profunda. Palabras Claves:estimación del esfuerzo de software, estimación del coste de software, regresión mediante procesos gaussianos, árbol de modelos, stacking, benchmark, MMRE, PRED(25), conjunto de datos similar a COCOMO, clasificación de regresores Smart citations: SciteAI. Dimensions.Open Alex.

Read PDF

Similar papers

Review

Evaluation of Software Effort Estimation Methods for Machine learning techniques

A review of deferent machine learning methods that are using for effort estimation like regression models, decision trees, random forest, neural networks, and then evaluate this models based on performance criteria such as MAE (Mean Absolute Error) and R 2 Score.

Montaser Fadulalla Ahmed Adam, Haroun Abdalla Eissa · 0 citations
Open access Aug 2026

Improving Software Effort Estimation Through Feature Selection and Optimized SVR

Accurate software development effort estimation is essential but often hindered by high-dimensional data and the inefficiencies of handling feature selection and parameter tuning as separate, sequential processes. This study proposes an integrated Whale Optimization Algorithm–Support Vector Regression (WOA-SVR) framewo...

R. Putri, G. E. Yuliastuti, Citra Nurina Prabiantissa · 0 citations
Open access Aug 2026

An Explainable Feature Selection and Stacking Ensemble Framework for Software Fault Prediction

Overall, the findings indicate that integrating principled feature selection with a boosting-based stacking ensemble can improve software fault prediction performance while providing greater transparency for software quality management.

Harsimran Kaur, Hardeep Singh, Amitpal Singh Sohal et al. · 0 citations
Open access Aug 2026

Comparative Evaluation of TabKANet with Oversampling and Feature Selection Ablation for Software Defect Prediction

TabKANet is a competitive architecture for all-numerical, highly imbalanced SDP, matching strong neural baselines and surpassing TabNet, where effective class weighting alone suffices and SMOTE is counter-productive.

Muhammad Faza Azhiman Saputra, Setyo Wahyu Saputro, M. Faisal et al. · 0 citations
Open access Aug 2026

A Robust Heterogeneous Ensemble Framework for Software Cost Estimation with Prediction Uncertainty Quantification

This paper presents a method using an uncertainty-aware heterogeneous ensemble of 100 bootstrap-trained base learners, which jointly produce point predictions and prediction intervals using a robust trimmed mean, with effort modeled on the log scale.

H. Sharif, Tara Nawzad Ahmad Al Attar, D. Rashid · 0 citations

TabKANet for Software Defect Prediction: An Oversampling and Feature-Selection Ablation Study

This study adapts TabKANet to the all-numerical, highly imbalanced SDP setting and empirically evaluates it against established baselines, using a structured ablation in order to isolate the contribution of oversampling and feature selection rather than to propose a new architecture.

Setyo Wahyu Saputro, M. Faza, Azhiman Saputra Setyo et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.