Comparative Evaluation of LSTM-Based Deep Learning Models for Software Effort Estimation
Abstract
Software development effort estimation is a critical aspect of effective project planning, as inaccurate predictions can lead to cost overruns, schedule delays, and incomplete system implementation. This study evaluates five LSTM-based deep learning architectures—Standard LSTM, CNN-BiLSTM, Residual LSTM, LSTM-GRU, and LSTM-Transformer—to determine the most robust model across diverse project conditions. The architectures (Standard LSTM, CNN-BiLSTM, Residual LSTM, LSTM-GRU, and LSTM-Transformer) are tested on three public datasets (NASA, UCP, and NASA93). Performance is measured using Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE), while the Wilcoxon Signed-Rank Test assesses statistical significance. Results indicate that LSTM-GRU consistently delivers stable performance across datasets, while LSTM-Transformer achieves superior accuracy on complex datasets due to its self-attention mechanism. This work offers a comprehensive, statistically rigorous comparison of five LSTM-based hybrid architectures, highlighting their relative strengths and weaknesses, and provides practical guidance for selecting deep learning models for reliable, data-driven software effort estimation.