Skip to content
Open access

Reinforcement Learning for Resource Allocation in Energy-Harvesting Cooperative IoT Networks

Jul 2026 · Network · Vol 6, pp. 49 · 0 citations · 36 references
Computer Science

TL;DR

This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.

Abstract

The internet of things (IoT) has rapidly evolved into a ubiquitous communication paradigm for enabling the deployment of autonomous wireless networks across diverse application domains. However, the limited energy storage capacity and computational resources of IoT devices (IoTDs) pose a serious concern to their long-term sustainability and the expected quality of service delivery. Moreover, in the foreseeable era of the internet of everything, centralised network resource management is likely to constrain network scalability. To tackle these challenges in the current and next-generation communication networks, the adoption of adaptive and lightweight computational frameworks coupled with energy-efficient transmission strategies is essential. To demonstrate this, we exploit the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs. Furthermore, to intelligently and autonomously perform resource allocation, we employ the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning. Based on key performance evaluation metrics, we compare our findings with the baseline methods, including the equal, random, and greedy power level selection schemes, with SARSA exhibiting the most favourable performance trade-offs.

Read PDF

Similar papers

2026

Towards Sustainable IoT: An AI-Driven Framework for Enhanced Energy Harvesting in Wireless Sensor Networks

An AI-driven framework integrates hybrid energy harvesting mechanisms with Deep Reinforcement Learning (DRL) to optimize energy efficiency in IoT systems and achieves up to 300% improvement in network lifetime under low-energy harvesting conditions.

Elkhatim Abuelysar Elmobarak Mohammed Ali · 0 citations
Open access Jul 2026

RLIOT: REINFORCEMENT LEARNING - BASED NETWORK RESOURCE OPTIMIZATION USING IOT SENSOR DATA

Results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.

L. Hoang, Van-Tam Hoang, Huu-Huy Ngo · 1 citation
Conference Jul 2026

DEERL-WSN: A Distributed Energy-Efficient Reinforcement Learning Approach for Wireless Sensor Networks in IoT Applications

Wireless Sensor Networks (WSNs) play a crucial role in the expanding landscape of the Internet of Things (IoT), yet they continue to face persistent challenges related to energy consumption, computational efficiency, and scalability. Although protocols like the Energy-Efficient Routing Protocol through Hybrid Algorithms (EERHA) have made progress in extending network lifespan using machine learning, they still encounter limitations particularly in managing processing overhead, adapting to changing conditions, and scaling to larger deployments. Unlike existing approaches that merely combine machine learning modules, DEERL-WSN (Distributed Energy-Efficient Reinforcement Learning for Wireless Sensor Networks), introduces a unified hierarchical learning framework where distributed reinforcement learning, lightweight graph neural networks, transfer learning, and federated optimization operate cooperatively. The novelty lies in (i) adaptive reward-driven routing using meta-learned objective weights, (ii) topology-aware clustering through lightweight GNN embeddings with significantly reduced computational complexity, (iii) transfer learning-assisted cluster-head prediction to eliminate repetitive optimization overhead, and (iv) federated deep reinforcement learning enabling scalable learning without centralized processing bottlenecks. These integrated innovations collectively address energy efficiency, scalability, and computational constraints simultaneously, which remain largely unresolved in existing WSN routing frameworks. DEERL-WSN addresses several bottlenecks found in earlier protocols and the simulation results show that DEERL-WSN significantly outperforms EERHA and other state-of-the-art methods.

Maheshkumar Patil, B. J, K. R et al. · 0 citations
Open access Jul 2026

A dynamic reward framework for scalable and efficient IoT-WSN routing using deep reinforcement learning.

Internet of Things-based wireless sensor networks (IoT-WSNs) face persistent challenges related to energy consumption, latency, and network congestion under dynamic and heterogeneous topologies. Conventional reinforcement learning approaches rely on static reward formulations, which limit adaptability and hinder effective multi-objective optimization. This study proposes a dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs. The proposed approach employs real-time reward recalibration to jointly optimize energy efficiency, delay, and throughput under varying network conditions. A hybrid deep reinforcement learning architecture is developed by integrating value-based, policy-based, and actor-critic methods, along with multi-agent coordination and attention mechanisms to prioritize critical nodes and links. Furthermore, a hierarchical learning structure decomposes global and local routing objectives, improving scalability and decision efficiency in complex network environments. Experimental results demonstrate that the proposed framework achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods. These findings highlight the effectiveness of dynamic reward adaptation for scalable and robust multi-objective optimization in IoT-WSN routing.

Suresh Betam, S. Nagendram, Bathula Prasanna Kumar et al. · 0 citations
Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations