Comparative Evaluation of LSTM and BiLSTM for Google Stock Price Forecasting
Abstract
Unidirectional Long Short-Term Memory (LSTM) and Bidirectional LSTM (BiLSTM) models are types of Recurrent Neural Networks (RNNs) capable of modelling long-range dependencies in financial time series, such as stock prices. However, the performance of LSTM and BiLSTM is highly sensitive to hyperparameter tuning, and there is limited guidance on systematic tuning, particularly for nonlinear stock-price series with varying historical lengths. Hence, this study aims to optimise the forecasting performance of LSTM and BiLSTM by tuning the hyperparameters through grid search and examining their behaviour. Daily Google stock prices were collected from Yahoo Finance, covering 3 January 2005 to 25 May 2022, and the adjusted closing price was selected as the target because it reflects corporate actions and offers more stable price movement. Two datasets were analysed, Sample 1 (25 May 2017–25 May 2022) and Sample 2 (3 January 2005–25 May 2022). Each dataset was split using an 80:20 train–test design. Min–Max normalisation was fitted on the training set and applied to the test set using the same parameters. Both architectures were implemented with a single recurrent layer and a single output neuron for regression, trained using the Adam optimiser and Mean Squared Error (MSE) loss. A grid search considered hidden neurons (32, 64, 128), epochs (20, 50, 100, 150), and batch sizes (32, 64, 128). The results show that in Sample 1, the best LSTM achieved MSE 8.4748 and MAE 2.2192, while the best BiLSTM achieved MSE 7.4766 and MAE 2.1031, both at 128 neurons, 150 epochs and batch size 32. For Sample 2, the best LSTM achieved MSE 4.0303 and MAE 1.4567, and the best BiLSTM achieved MSE 3.9929 and MAE 1.4373 under the same hyperparameters. In conclusion, increasing the number of hidden neurons and training more epochs generally reduced errors, while larger batch sizes tended to increase errors, and BiLSTM yielded lower average MSE and MAE than LSTM across the two samples.