A Multimodal Hybrid LSTM–CNN Framework for Robust Spatio-Temporal Air Quality Index Prediction
Abstract
Air pollution is a severe challenge to human health and urban sustainability, and accurate and timely prediction of the Air Quality Index (AQI) is critical to early warning systems and policy interventions. Current statistical and single-modality deep learning models are ineffective in capturing the non-linear spatio-temporal dynamics of atmospheric pollutants. This paper proposes a multimodal hybrid deep learning architecture, which combines Long Short-Term Memory (LSTM) networks to model the temporal sequence of pollutants and Convolutional Neural Networks (CNNs) to extract spatial features from environmental images. This approach presents an innovative strategy for cross-modal feature-level fusion of multimodal inputs, distinguishing it from traditional hybrid LSTM–CNN networks. The model uses a Softmax output layer to classify six categories of AQI (Good, Moderate, Unhealthy to Sensitive Groups, Unhealthy, Very Unhealthy, and Hazardous). The hybrid model performs better than baseline LSTM, CNN, and fusion models in terms of classification accuracy (94%), F1-score (0.957), and reduced AQI prediction errors, with Mean Absolute Error (MAE) = 7.0 and Root Mean Squared Error (RMSE) = 9.0, on the Air Pollution Image Dataset of India and Nepal (12,000+ annotated samples). These findings show that multimodal spatio-temporal learning is very effective in predicting AQI with high reliability and has strong potential for real-time smart-city air quality monitoring and decision-support systems.