Aug 2026· مجلة جامعة صنعاء للعلوم التطبيقية والتكنولوجيا· Vol 4, pp. 2416-2427· 0 citations· 24 references
TL;DR
Examination of aggregation stability and feasibility-sensitive aggregation in FRL for Edge-IoT systems identifies a need for aggregation mechanisms that jointly account for update stability, update reliability, resource availability, and constraint feasibility.
Abstract
Federated Reinforcement Learning (FRL) provides a useful basis for distributed policy learning in Edge-IoT systems, where clients interact with local environments without transferring raw operational data to a central server. Yet aggregation becomes difficult when clients operate under different transition dynamics, workloads, resource capacities, communication conditions, and operational constraints. In these settings, local policy updates may not differ only in magnitude or direction; they may also differ in stability, reliability, resource support, and operational feasibility. Conventional averaging is therefore limited, since it does not distinguish stable and feasible updates from unstable or constraint-violating ones. This paper examines aggregation stability and feasibility-sensitive aggregation in FRL for Edge-IoT systems. It reviews and synthesizes related literature across four connected streams: heterogeneous Federated Learning, FRL-based edge decision-making, constrained and safe Reinforcement Learning, and adaptive or reliability-aware aggregation. The reviewed studies are analyzed through six dimensions: learning paradigm, type of heterogeneity, role of policy learning, treatment of operational constraints, aggregation strategy, and whether local feasibility signals influence global aggregation weights. The analysis indicates that existing studies provide valuable foundations, but they usually treat heterogeneity, constraint handling, and aggregation adaptation as separate concerns. The paper identifies a need for aggregation mechanisms that jointly account for update stability, update reliability, resource availability, and constraint feasibility. It positions aggregation as adaptive client influence regulation rather than passive averaging in future Edge-IoT FRL systems.
A federated reinforcement learning framework for adaptive load balancing in the edge-fog-cloud continuum that optimizes energy efficiency and supports diverse quality of service requirements and uses a distributed experience replay buffer to reduce trial-and-error in reinforcement learning.
Si Liu, Midhun Chakkaravarthy· Future Technology· 0 citations
The Federated Green Anaconda Optimizer (FedGAO), an innovative FL framework inspired by the behavioral patterns of the Green Anaconda Optimizer (GAO), is proposed, demonstrating superior performance in terms of accuracy, convergence speed, and resource efficiency.
Elahe Eslami, S. A. Shahzadeh Fazeli, J. Abouei et al.· Cluster Computing· 0 citations
Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation, higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines.
Jayesh Kumar Dabi, Priyadarshi Ashok Dahat· International Journal of Wir...· 0 citations
Industrial IoT predictive maintenance demands real-time anomaly detection under tight resource and interpretability constraints, while monolithic LLM-based systems remain impractical for on-site deployment. We introduce HAMA (Hierarchical Adaptive Multi-Agent Architecture), in which “adaptive” refers strictly to online...
Rebin Saleh, K. Dinh, Balázs Villányi et al.· IEEE Access· 0 citations
The rapid growth of large-scale interconnected systems, such as smart cities, industrial automation, and environmental monitoring, demands intelligent decision-making frameworks that are resilient, scalable, and resource-efficient. Traditional centralized intelligence approaches suffer from communication bottlenecks, h...
M. Kishore, N. Velmurugan· 2026 7th International Confe...· 0 citations
A packet-level transmission framework that captures buffer overflow, delay violations, and transmission errors, and uses the resulting packet delivery ratio (PDR) to represent partial-update reception through a packetized, Bernoulli-masked FL aggregation process is developed.