Intelligent system for analysis and assessment of population health harm risks by air pollution level
Abstract
Air pollution by fine particulate matter (PM2.5 , PM10 ), nitrogen dioxide (NO2 ), sulfur dioxide (SO2 ), carbon monoxide (CO) and ozone (O3 ) is among the leading environmental threats to public health in densely populated urban areas, yet the tools available to health authorities rarely combine heterogeneous monitoring data with interpretable risk estimates. This paper presents an intelligent system that assigns an observation to one of six air quality index (AQI) classes, used as a proxy for the level of harm to population health, together with a comparative evaluation of the six models behind it: a decision tree, a support vector machine, a random forest, XGBoost, a convolutional neural network (CNN) and a long short-term memory (LSTM) network. All six are trained on identically preprocessed data – the Air Pollution Image Dataset from India and Nepal, cross-checked with structured measurements from two open sources – and are compared at the evaluation stage rather than chained into a sequence. The system is implemented as a client-server application with a Python and Flask backend and a React frontend, whose dashboard provides exploratory data analysis, model inference and model comparison. The random forest and the CNN are the most accurate models (accuracy 0.988 and 0.987, macro-averaged F1-score 0.987 and 0.986), followed by the LSTM network (0.980 and 0.979), whereas the single decision tree falls to 0.856 because it does not separate the two most severe classes. Average training time ranges from 3 s for the decision tree to 67 s for the LSTM network, which makes the random forest the best compromise between accuracy and computational cost.