Performance and Limitations of Machine Learning Models for Groundwater Level Prediction Using Hydro-Climatic Variables
Accurate groundwater level (GWL) prediction is essential for sustainable groundwater management and resource planning. However, it is challenging in heterogeneous hydrogeological settings for physics-based models, particularly under limited subsurface characterisation. Machine learning (ML) techniques can capture complex spatio-temporal groundwater dynamics, complementing conventional modelling approaches. This study systematically evaluates the performance and limitations of ML models for GWL prediction in sandstone and mudstone formations, with emphasis on the influence of local hydrogeological conditions. Artificial Neural Network (ANN), Long Short-Term Memory (LSTM), Support Vector Regression (SVR), and Random Forest (RF), along with their wavelet-enhanced counterparts, were applied to monthly hydro-climatic data from 11 observation wells in the Lower Otter Catchment, UK, covering 2011–2023. The final four years were reserved for validation. Time-series predictors were used because time-invariant and sparsely available geological parameters provide limited explanatory power for local-scale GWL dynamics. Their influence is implicitly reflected in observed GWL responses. Model performance was assessed using statistical criteria, including the coefficient of determination (R 2 ). Additionally, a new metric, the Data Difference and Trend Index (DDTI), was introduced to measure the proportion of simulated values matching observed trends within a predefined threshold (e.g. 0.5 m). Model performance was site-specific, with validation R 2 ranging from < 0.1 to > 0.9. This indicates the dominant influence of local hydrogeological conditions, with lower accuracy observed in partially confined, non-recharge-dominated, and river-disconnected wells. A 2-month time lag produced optimal model performance, reflecting the catchment’s characteristic response time to infiltration processes. Individual models occasionally outperformed ensemble averages, which showed fewer outliers. Wavelet transforms did not consistently enhance performance. Model efficacy varied seasonally, with validation R 2 markedly lower in summer (e.g. < 0.1) and higher in autumn (e.g. > 0.9). This emphasises key limitations of ML-based GWL prediction, including reduced reliability near lithological boundaries and strong sensitivity to hydro-climatic conditions, constraining model transferability. Overall, the findings highlight the value of moving beyond performance benchmarking to explicitly identify hydro-climatic and hydrogeological conditions under which ML models lose reliability, informing groundwater modelling and sustainable water management. Graphical Abstract This study evaluates the performance and limitations of machine learning (ML) models for predicting groundwater levels (GWL) in sandstone and mudstone formations using hydro-climatic variables across 11 observation wells in the Lower River Otter Water Body, UK. Four ML models, i.e. Artificial Neural Networks (ANN), Long Short-Term Memory (LSTM), Support Vector Regression (SVR), and Random Forest (RF), along with their wavelet-enhanced versions, were applied to monthly hydro-climatic data from 2011 to 2023, with the last four years reserved for validation. Model performance was assessed using the coefficient of determination (R 2 ), Root Mean Square Error (RMSE), and Nash-Sutcliffe Efficiency (NSE), and a novel Data Difference and Trend Index (DDTI), which quantifies the proportion of simulated data following observed trends within a defined threshold. Results indicate that predictive accuracy is highly site-specific, with local hydrogeological conditions strongly influencing outcomes. No model consistently captured GWL near interior boundaries where sandstone is confined by mudstone, and neither wavelet transforms nor model ensembles reliably improved performance. Seasonal variability also affected model efficacy, with the highest accuracy in autumn and the lowest in summer. Overall, the workflow highlights the limitations of ML for GWL prediction and provides insights for future hydrogeological modelling.