Jul 2026· International journal of Computer Networks & Communications· Vol 18, pp. 01-16· 1 citation· 22 references
TL;DR
Results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.
Abstract
The fast rise of wireless communication networks, including 6G, Internet of Things (IoT), and edge com puting, has created unprecedented demand for spectrum and energy resources.become a significant challenge in modern IoT networks due to heterogeneous devices, dynamic traffic patterns, and diverse QoS requirements. This study proposes a Deep Reinforcement Learning (DRL)–basedframework for optimizing network resource allocation in IoT environments using real-world sensor data. The proposed framework differs from existing studies that typically assess reinforcement learning methodologies under simplified wireless network assumptions and idealized conditions. Our method functions on heterogeneous IoT traffic produced by various device types, including sensors, actuators, and cameras, each possessing distinct Quality of Service (QoS) requirements. To ensure practical applicability, a realistic IoT simulation environment is developed, incorporating dynamic bandwidth release and queue-aware resource management to emulate real-world network behavior. Furthermore, a Deep Q-Network (DQN) agent with an enhanced exploration strategy is designed to improve learning stability and convergence performance, enabling more efficient and adaptive resource allocation in dynamic IoT scenarios. Experimental results show that the proposed DQN agent achieves a 26.7% improvement in cumulative reward compared to a random policy and consistently outperforms conventional heuristic approaches. This significant gain indicates that the agent effectively learns a structured resource allocation strategy rather than making uninformed decisions. These results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.
This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.
Olumide Alamu, T. Olwal, Emmanuel M. Migabo· Network· 0 citations
Internet of Things-based wireless sensor networks (IoT-WSNs) face persistent challenges related to energy consumption, latency, and network congestion under dynamic and heterogeneous topologies. Conventional reinforcement learning approaches rely on static reward formulations, which limit adaptability and hinder effective multi-objective optimization. This study proposes a dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs. The proposed approach employs real-time reward recalibration to jointly optimize energy efficiency, delay, and throughput under varying network conditions. A hybrid deep reinforcement learning architecture is developed by integrating value-based, policy-based, and actor-critic methods, along with multi-agent coordination and attention mechanisms to prioritize critical nodes and links. Furthermore, a hierarchical learning structure decomposes global and local routing objectives, improving scalability and decision efficiency in complex network environments. Experimental results demonstrate that the proposed framework achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods. These findings highlight the effectiveness of dynamic reward adaptation for scalable and robust multi-objective optimization in IoT-WSN routing.
Suresh Betam, S. Nagendram, Bathula Prasanna Kumar et al.· Scientific Reports· 0 citations
Smart kitchen IoT environments integrate robotic manipulators, sensing modules, and intelligent appliances within dense wireless deployments. Such environments generate heterogeneous traffic including latency-critical control signals, high-bandwidth multimedia streams, and best-effort telemetry. Conventional MQTT deployments rely on static QoS-based acknowledgment (ACK) behaviors that cannot adapt to dynamic congestion and packet loss. This paper proposes RL-ACK, a reinforcement learning–based adaptive acknowledgment framework implemented at the edge router. Per-message ACK selection is formulated as a Markov Decision Process (MDP), and the ACK policy is learned using a Deep Q-Network (DQN). The reward function balances latency, control-plane overhead, reliability violations, and utilization pressure. Extensive NS-3 simulations and a real Wi-Fi testbed demonstrate up to 36.7% latency reduction, approximately 22% signaling overhead reduction, and ACK-induced energy reduction for Tier 3 (BET) devices compared with static MQTT QoS 1, while maintaining Tier 1 deadline-compliant delivery above 95% under high network utilization ( $\rho = 0.9$ ).
Jinsuk Baek, Sanho Lee, Ju Hong Park· IEEE Access· 0 citations
Robust Reinforcement Learning (RL) based task scheduling approaches can address the inherent tradeoff between energy consumption and deadline violation in a Multi-access Edge Computing (MEC) based Internet of Things (IoT) network, while maintaining robustness against changes in the task arrival rate. However, tabular robust RL algorithms suffer from high computational and storage complexity, and therefore are not scalable to a system with a large number of IoT nodes. To this end, in this paper, we propose a robust deep RL based task scheduling algorithm to solve the underlying Robust-Return Constrained Markov Decision Process (R2CMDP) problem. The proposed algorithm introduces a tunable amount of robustness in the solution of the RL framework. Complexity analysis and ns-3 simulation results are presented to demonstrate the efficacy of our algorithm.
V. Masih, Arghyadip Roy· International Conference on...· 0 citations
Wireless Sensor Networks (WSNs) play a crucial role in the expanding landscape of the Internet of Things (IoT), yet they continue to face persistent challenges related to energy consumption, computational efficiency, and scalability. Although protocols like the Energy-Efficient Routing Protocol through Hybrid Algorithms (EERHA) have made progress in extending network lifespan using machine learning, they still encounter limitations particularly in managing processing overhead, adapting to changing conditions, and scaling to larger deployments. Unlike existing approaches that merely combine machine learning modules, DEERL-WSN (Distributed Energy-Efficient Reinforcement Learning for Wireless Sensor Networks), introduces a unified hierarchical learning framework where distributed reinforcement learning, lightweight graph neural networks, transfer learning, and federated optimization operate cooperatively. The novelty lies in (i) adaptive reward-driven routing using meta-learned objective weights, (ii) topology-aware clustering through lightweight GNN embeddings with significantly reduced computational complexity, (iii) transfer learning-assisted cluster-head prediction to eliminate repetitive optimization overhead, and (iv) federated deep reinforcement learning enabling scalable learning without centralized processing bottlenecks. These integrated innovations collectively address energy efficiency, scalability, and computational constraints simultaneously, which remain largely unresolved in existing WSN routing frameworks. DEERL-WSN addresses several bottlenecks found in earlier protocols and the simulation results show that DEERL-WSN significantly outperforms EERHA and other state-of-the-art methods.
Maheshkumar Patil, B. J, K. R et al.· 2026 7th International Confe...· 0 citations