A Federated Predictive Predictive Load Balancing (FPLB) framework is proposed to combine Long Short-Term Memory (LSTM) workload forecasting with federated learning, which does not require fog nodes to share their operational data.
Abstract
While the advantages of fog computing in delivering low-latency Internet of Things (IoT) applications are well understood, efficient load balancing remains a constant challenge due to the diversity of node capabilities and the uncertainty of workloads. Current scheduling methods are typically reactive and only take action once they detect congestion, and require gathering data centrally, which can be privacy and bandwidth sensitive. In this paper, a Federated Predictive Load Balancing (FPLB) framework is proposed to combine Long Short-Term Memory (LSTM) workload forecasting with federated learning, which does not require fog nodes to share their operational data. Predicted workloads feed a normalized load index for proactive task assignment, while a differential-privacy mechanism with a Rényi accountant protects model updates during federated aggregation. All experiments are reported from a self-contained simulator. Across eight independent seeds under a moderate-to-high load, FPLB attains the lowest average task latency (149.2 ms), significantly below Deep Q-Network (DQN) scheduling (1.6% reduction; p < 0.01, Wilcoxon signed-rank) and a federated-DQN control, and well below reactive heuristics (24.3% below round-robin). The margin widens with load, reaching 2.9% over DQN at 10 tasks/s, and FPLB’s latency variance is consistently the lowest, indicating more predictable scheduling. An ablation confirms that workload prediction is the primary driver of the improvement and that federation lowers prediction error. Federated communication overhead is 0.4% of network traffic at 50 nodes, rising to only 3.7% at 500 nodes, and performance is robust to 30% per-round node dropout. This paper further characterizes the privacy-utility envelope as the budget tightens from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\:\epsilon\:$$\end{document} = 3.2 to 0.5. At low load, where congestion is rare, DQN is comparable, so the framework is most valuable for deployments that regularly experience dynamic or peak-heavy demand.
The growing use of cloud services in the contemporary society has put unprecedented pressure on scalability,
responsiveness, and fault resilience, especially in multi-cloud environments that combine heterogeneous resources across
providers. Such systems have been difficult to balance their loads effectively because o...
Damodhar MummiReddy· ˜The œinternational Arab jou...· 0 citations
An asynchronous framework named FedQS, which employs a multi-dimensional staleness evaluation mechanism that dynamically assesses updates by combining the similarity between local and global models with client latency metrics, and implements a decoupling solution via a queue scheduling algorithm to resolve the coupling...
Jia-Hui Zhou, Fang Li, Tian-Yu Shi et al.· Journal of Cloud Computing· 0 citations
Federated Learning (FL) enables collaborative model training across distributed devices. A major concern in FL is how to operate it smoothly in resource-constrained environments, where this process must perform under strict operational constraints, such as cost or energy budgets. An often-overlooked aspect is the effec...
Anna Lackinger, P. Frangoudis, Andrea Morichetta et al.· 2026 International Conferenc...· 0 citations
With the proliferation of Internet of Things (IoT) applications, a massive amount of data has been produced, requiring an efficient platform to store and process this data. Cloud computing has the ability to tackle such enormous data, but cannot provide real-time response to latency sensitive IoT applications. Fog comp...
M. Aknan, Maheshwari Prasad Singh, Rajeev Arya· International Journal of Int...· 0 citations
Cloud Computing has played a vital role in handling data storage, access and processing in distributed environments. Cloud Computing delivers scalable servers, databases and storage resources and abstracts the backend infrastructure to the end users. Workloads exhibit higher instability due to the increase in the depen...
Raghav Rs, Vignesh R, Arunprasad P et al.· 2026 6th International Confe...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026