Flood Hazard Assessment in Jabodetabek: Benchmarking Three Machine Learning Classification Models
Abstract
Flood disasters are Indonesia's most persistent and economically disruptive natural hazard, with the Jabodetabek metropolitan region particularly vulnerable because of its low-lying coastal topography, high population density, progressive land subsidence, and accelerating urbanization-driven imperviousness. Conventional hydrological assessment tools frequently underestimate the spatial heterogeneity of flood hazards in complex urban environments, motivating the development of robust, data-driven alternatives. This study presents a machine learning-based flood classification framework integrating meteorological parameters (average and maximum rainfall, temperature, and soil moisture), topographic characteristics (elevation and slope), and satellite-derived NDVI to quantify and spatially characterize flood hazard impacts. Three classification algorithms, such as Logistic Regression, Random Forest, and Gradient Boosting, were benchmarked on a balanced dataset of 3,000 georeferenced environmental records from the Jabodetabek area. The Gradient Boosting Classifier achieved the highest predictive performance: accuracy 94.00%, F1-Score 0.929, and ROC-AUC = 0.971. Feature importance analysis identified NDVI and average rainfall as the dominant predictors, while temperature exhibited a strong negative correlation (ρ = -0.56), reflecting the thermodynamic signature of monsoon convective systems. These findings provide a reproducible framework for evidence-based hazard zonation and disaster risk management in Indonesian urban contexts.