MACHINE LEARNING-BASED PREDICTION OF WATER QUALITY USING PHYSICOCHEMICAL AND ENVIRONMENTAL INDICATORS
Abstract
Reliable prediction of water quality can strengthen environmental monitoring by integrating conventional measurements with data-driven analytical approaches. This study evaluated the capacity of physicochemical, environmental, spatial, and hydrological indicators to predict an integrated water-quality score using machine-learning regression. The analysis included 36 complete observations from Anuppur, Dindori, and Jabalpur and incorporated indicators of oxygen conditions, organic loading, nutrients, ionic composition, microbiological contamination, rainfall, water level, season, and location. Four algorithms—Ridge Regression, Random Forest, Gradient Boosting, and Decision Tree Regression—were evaluated using six-fold cross-validation. Distinct spatial patterns were observed, with Jabalpur recording the highest mean water-quality impairment score and Anuppur the lowest. Total coliform, phosphate, BOD, conductivity, dissolved solids, and COD showed strong positive associations with the integrated outcome. Ridge Regression achieved the strongest overall predictive performance (R² = 0.603; RMSE = 6.98; MAE = 5.01), while Random Forest produced the lowest MAE (4.93). The findings demonstrate that a parsimonious machine-learning framework can capture meaningful water-quality variability from multidimensional environmental measurements. Although the limited sample size constrains generalizability, the approach provides a practical basis for identifying influential indicators and supporting targeted water-quality monitoring.