Refined simulation of groundwater level dynamics based on a machine learning stacking framework
Abstract
Accurate and rapid prediction of groundwater levels (GWL) is essential for effective groundwater management. Machine learning models are efficient tools for GWL prediction, but individual models often suffer from limited generalization due to inherent randomness. This study proposed a stacking-based GWL prediction framework suitable for arid regions in Northwest China. Feature variables affecting GWL were selected using variable importance in projection (VIP). Then, three machine learning models—artificial neural networks (ANN), random forests (RF), and Light Gradient Boosting Machine (LightGBM)—were developed, and their outputs were integrated using a support vector regression (SVR)-based stacking method to enhance the accuracy of GWL prediction. The results show that the factors of influencing GWL changes vary significantly across different regions, and selecting the most contributive feature variables is beneficial for model construction. Among the individual models, the RF model demonstrated higher accuracy and more stable performance, outperforming the ANN and LightGBM models. However, individual models exhibited poor generalization during validation. In contrast, the stacking model maintained high performance, demonstrating superior generalization. Compared to the best-performing individual model (RF) in validation period, the Nash–Sutcliffe efficiency ( NSE ) and Kling–Gupta efficiency ( KGE ) of stacking model improved by 0.11–0.66 and 0.05–0.41, the correlation coefficient ( R 2 ) increased by 0.05–0.3, and root mean square error ( RMSE ) reduced by 0.01–0.1 m. In the stacking simulation, RF had the highest average contribution (80.2%), followed by ANN (13.9%) and LightGBM (5.9%). This study provides a stacking simulation framework based on machine learning methods for precise groundwater level simulation, which can serve as a reference for groundwater level simulation in other regions.