Skip to content
Open access

The Impact of Data Preprocessing on Federated Learning for Air Quality Prediction: A Comparative Study of Imputation Methods

Aug 2026 · Mathematical Modeling and Algorithm Application · 0 citations · 15 references

TL;DR

It is shown that preserving data quality should take priority over quantity when the missing rates are exceptionally high, providing the practical guidance for robust preprocessing design in federated learning systems.

Abstract

Data preprocessing plays a foundational role in machine learning, but it receives limited systematic attention in federated learning (FL) environments. This study empirically compares four missing value strategies using the UCI dataset: direct deletion, mean imputation, K-Nearest Neighbors (KNN) imputation, and Random Forest (RF) imputation. The setup uses a federated framework with five clients, each running a multi layer perceptron model. The findings show direct deletion achieves the strongest performance, with a mean absolute error of 0.173 and an  of 0.966, clearly exceeding imputation methods (the errors around 0.25). The NMHC(GT) feature has an 88.4% missing rate, making imputation unreliable and introducing noise. Although direct deletion reduces the sample from 7,674 to 827 observations, it safeguards data integrity. This research shows that preserving data quality should take priority over quantity when the missing rates are exceptionally high, providing the practical guidance for robust preprocessing design in federated learning systems.

Read PDF

Similar papers

Sep 2026

Overcoming data shortages through hybrid machine learning in air quality forecasting: a case study of Kağıthane, Istanbul

A hybrid machine-learning framework for AQI prediction in Kağıthane, Istanbul, Türkiye, using a long-term dataset spanning approximately ten years, thereby reducing reliance on dense sensor infrastructures and achieving overall accuracy of approximately 98%, enabling timely health advisories and more efficient allocati...

M. Akiner, M. Ghasri · 1 citation
Jul 2026

Management of missing air pollution data within urban environments using machine learning regressions: a case study for Delhi, India

An iterative, multi-model machine learning workflow to reconstruct missing daily pollution data for all six pollutants across 45 stations from 2014–2024 supports the usefulness of the approach for long-term regional reconstruction while also highlighting its limitations for pollutants with strong local emission signatu...

Sedra Shafi, Nicola Scafetta · 0 citations
Review Open access Aug 2026

An Investigation of Federated Learning for Air Quality Forecasting and Monitoring

There are significant gaps that remain in terms of model interpretability and the ability to generalize across climate variations, so this article provides a relatively comprehensive overview of the application of federated learning in air quality forecasting and monitoring.

Yuhao Wu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.