Skip to content
Open access

A Machine Learning Framework for Flood Risk Assessment With Imbalanced Data

Sep 2026 · Journal of Flood Risk Management · 0 citations · 34 references

Abstract

This study presents a machine‐learning framework for identifying areas of elevated flood risk using imbalanced, high‐dimensional geospatial datasets. Using the Mandra River Basin in Greece as a case study, an extreme gradient boosting (XGBoost) classifier was trained on 11 flood‐related features including hydrological, meteorological, soil and topographic data, whereas 454 geolocated citizens' emergency calls were used as target information. Class imbalance was addressed through a weighted‐only strategy to preserve the integrity of the original dataset, while adjusting class contributions during the training process. Bayesian optimization and stratified 10‐fold cross‐validation were combined to tune and evaluate the model performance. Threshold‐sweep analysis enables the identification of multiple operational settings based on precision–recall trade‐offs and cost‐sensitive criteria. The results revealed physically interpretable feature importance patterns, with terrain elevation (TE) (45%) and accumulated rainfall (25%) emerging as the dominant predictors according to feature importance analysis, followed by curve number (16%) and the Manning coefficient (9.2%). The proposed framework provides a reliable, data‐driven tool for flood risk assessment, and supports a flexible, two‐tier decision strategy that differentiates between emergency escalation and broader situational awareness.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.