Enhanced Air Quality Index Forecasting through Genetic Algorithm-Based Kernel Extreme Learning Machine
Abstract
Air quality is one of the significant issues that affect public health and environmental sustainability as well as urban development. It is important to make accurate forecasts of the air quality index to enable the stakeholders to predict the pollution levels and put appropriate mitigation measures in place. However, traditional machine learning methods have shown to have low prediction accuracy when used to predict air quality due to the complexity of the data. In an attempt to improve AQI prediction accuracy, this paper aims at developing a Genetic Algorithm-based Improved Kernel Extreme Learning Machine (GA-KELM). The suggested method introduces a kernel to the traditional Extreme Learning Machine algorithm in order to replace the hidden layer output matrix, thereby improving the non-linear learning capabilities and generalization of the model. In addition, the genetic algorithm is used to tune several model parameters such as number of hidden neurons, network structure, and random initialization of input weights and bias threshold, which usually limits the prediction capability of traditional ELM algorithms. Afterward, the optimal output weights are estimated based on the least squares approach, and an accurate and efficient predictive model is derived. The proposed GA-KELM forecasting model is investigated through experiments using long-term data sets recorded by monitoring air pollution of a metropolitan city in China. The model predicts the concentrations of major atmospheric contaminants such as SO2, NO2, PM10, CO, O3, and PM2.5 as well as the AQI values. It is benchmarked against other popular forecasting models, such as the Community Multiscale Air Quality (CMAQ) Model, Support Vector Machine (SVM), and Deep Belief Network-Backpropagation (DBN-BP) Neural Network. Experimental results show that the proposed GA-KELM forecasting model exhibits higher prediction accuracy, smaller forecasting error, and better robustness compared to the existing models.