Skip to content
Open access

A Comprehensive Comparative Analysis of Machine Learning Models for Daily Precipitation Forecasting Using Satellite-Based Meteorological Data

Sep 2026 · Osmaniye Korkut Ata Üniversitesi Fen Bilimleri Enstitüsü Dergisi · 0 citations · 17 references

Abstract

Accurate daily precipitation forecasting is a significant and persistent challenge in hydrology and atmospheric sciences, pivotal for effective water resource management, agricultural planning, and extreme event risk assessment. The primary novelty of this study lies in its comprehensive and systematic comparative framework, designed to rigorously evaluate a wide spectrum of twelve distinct machine learning philosophies—ranging from classical statistics and baselines to advanced tree-based ensembles, neural networks, and hybrid techniques—on a rich, multivariate, long-term satellite-based dataset. A 24-year (2001-2024) daily time-series dataset from the NASA POWER project was utilized for a specific location. The study's framework is tripartite, integrating: (1) a statistical analysis of historical trends and extreme event frequencies, (2) a data-driven automated feature selection process to identify the most salient atmospheric predictors, and (3) a definitive comparative evaluation of the forecasting models using both regression and categorical metrics. The historical analysis, conducted on annual maximum precipitation using the Mann-Kendall test, revealed no statistically significant trend. The frequency analysis using the Gumbel distribution estimated the 100-year design rainfall to be approximately 129.82 mm/day, with a 95% confidence interval of [84.82, 146.83] mm/day. In the comprehensive model competition, the Artificial Neural Network (ANN) model achieved the highest performance with a coefficient of determination (R2) of 0.680. This was closely followed by high-performing tree-based ensemble models, namely XGBoost (R2: 0.674) and Random Forest (R2: 0.665), demonstrating the overall superiority of modern machine learning approaches for this task. A key and novel finding of this work is the empirically demonstrated unsuitability of sequential deep learning models like LSTM and GRU for this specific tabular data problem, evidenced by their markedly poor R2 scores (0.063 and 0.056, respectively). This research establishes that with robust, automated feature engineering, modern machine learning models can explain approximately 68% of the variance in daily precipitation, providing a valuable and rigorously validated tool for a wide range of hydrological applications and setting a realistic performance benchmark.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.