Skip to content
Open access

Intelligent Task Offloading and Energy Management for Battery-Less 6G Industrial IoT Using Multi-Agent Deep Reinforcement Learning

Aug 2026 · Sulaimani Journal for Engineering Sciences · Vol 12, pp. 254-267 · 0 citations

TL;DR

A constrained, risk-sensitive multi-agent reinforcement learning framework is presented for joint task offloading and EH scheduling in battery-less 6G industrial networks, yielding a Pareto-non-dominated, statistically validated policy.

Abstract

Battery-less Industrial Internet of Things (IIoT) devices that are energized by energy harvesting (EH) must both finish latency-critical tasks and avoid capacitor brownouts simultaneously, but prior multi-agent learning approaches only optimize average performance, penalizing energy safety as a soft constraint and thus providing no assurances for energy-neutral operation. In this paper, a constrained, risk-sensitive multi-agent reinforcement learning framework is presented for joint task offloading and EH scheduling in battery-less 6G industrial networks. Learned dual variables that constrain brownout and energy-neutrality as hard budgets augment a centralized-training, decentralized-execution MAPPO backbone, a conditional-value-at-risk (CVaR) distributional critic is used to down-weight worst-case brownout, and a 6G ambient-backscatter transmission mode is used to enable communication in energy-scarce regimes. The method, L-MAPPO-EH, has the highest return and task completion, and reduces the brownout rate by a factor of three, on average, over the best learning baseline, and backscatter gains improve completion by an additional twenty-six percentage points under rare EH regimes, yielding a Pareto-non-dominated, statistically validated policy.

Read PDF

Similar papers

Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations
Open access Jul 2026

An Energy-Efficient Multi-Agent Reinforcement Learning Approach for Spark Job Scheduling in Mobile Edge Computing

Comparative tests with PPO, FIFO, FAIR and HAS baselines confirm that multi-agent reinforcement learning can well capture the intrinsic scheduling patterns of complex mobile environments, providing an adaptive and energy-efficient scheduling solution for practical IoT deployments.

Haoyu Gu · 0 citations
Conference Jul 2026

A Personalized Federated Deep RL Framework for Carbon-Aware and Battery-Safe Smart Microgrid Intelligence in Edge-IoT Environments

The distribution of renewable energy resources and Edge-IoT infrastructures have brought new challenges in intelligent smart microgrid management, such as dynamic-energy-demand changes, carbon-heavy energy-scheduling, communication overhead, and battery degradation. The deployment of renewable energy resources, distributed battery storage systems, and Edge-IoT infrastructures has created a number of challenges in the management of smart microgrids, such as dynamic changes in energy demand, carbon-heavy energy-scheduling, communication overhead, and battery degradation. To tackle these challenges, this paper introduces a personalized Federated Deep Reinforcement Learning (GridMind-FDRL) framework for carbon-aware and battery-safe decentralized smart microgrid optimization. The proposed framework combines federated learning, deep reinforcement learning, edge intelligence, carbon-aware energy scheduling and battery-aware adaptive optimization with a centralized framework for energy management. Unlike traditional centralized optimization methods, GridMind-FDRL allows for collaborative learning among distributed microgrid nodes while maintaining privacy and enabling low latency real-time optimization in dynamic Edge-IoT environments. The framework was tested with different operating conditions of intermittent renewables, varying load levels, and battery stress conditions. The results of the experiment showed that the energy efficiency could be 93.86%, carbon reduction 24.36%, battery health preservation 90.42%, communication efficiency 88.08%, and decision latency reduction 34.21%. The acquired results support the scalability, sustainability, and smart energy optimization ability of the suggested framework for the next-generation decentralized smart energy ecosystems.

Jaichandran R, P. Marimuthu, K.Nethra Devi et al. · 0 citations
Conference Jul 2026

Joint AoI and SWIPT-Aware Scheduling via Multi- Agent Deep Reinforcement Learning

This work investigates the joint optimization of Age of Information (AoI) and energy harvesting (EH) in wireless edge computing systems, where edge servers not only process IoT data but also act as wireless power suppliers via simultaneous wireless information and power transfer (SWIPT). Building upon the asynchronous model-free fractional multi-agent reinforcement learning framework and the Lyapunov drift-plus-penalty (DPP) concept, we design a fractional-based reward function for AoI and construct a virtual queue to enforce long-term energy stability under battery storage constraints. The overall reward is formulated as a weighted sum, capturing the trade-off between timeliness and energy sustainability, with update decisions, task offloading, and power splitting ratios as key control variables. Simulation results demonstrate that the developed multi-agent deep reinforcement learning approach achieves superior AoI–energy trade-offs compared to related baseline algorithms. These findings highlight the effectiveness of our framework in balancing information freshness and sustainable energy harvesting under resource-constrained edge environments.

Kuang-Ting Liu, Jain-Shing Liu, Wan-Ling Chang · 0 citations
Open access Jul 2026

Toward Zero-Downtime Industrial IoT: Digital Twin-Enabled Predictive Wireless Power Transfer and Sensing Scheduling

Industrial Internet of Things (IIoT) networks require continuous, uninterrupted sensing operations despite the finite battery capacity of deployed IoT nodes. Conventional reactive energy management, where nodes switch to charging mode only after residual energy falls below a fixed threshold, cannot prevent depletion events and compromises network uptime. We propose a digital twin (DT)-enabled predictive scheduling framework in which a DT layer co-located with a multi-access edge computing (MEC) control center continuously mirrors the physical network state and generates H-slot look-ahead scheduling decisions before depletion can occur. The framework operates over a 5G network-sliced infrastructure with dedicated URLLC, eMBB, and mMTC slices. Two coupled integer programming problems are formulated, namely a predictive IoT node scheduling problem and a predictive energy transmitter scheduling problem. Optimal solutions are obtained via branch-and-bound with reliability branching (DT-PBB), and a low-complexity DT-Aware Greedy Priority Heuristic (DT-GPH) is also proposed. Evaluated against Earliest-Deadline-First (EDF-WPT), No-WPT (a baseline that disables wireless charging entirely), and Random baselines across three parameter configurations with K up to 200 nodes, DT-PBB achieves the highest sensing utility and the fewest energy depletion events in all scenarios. DT-GPH provides near-optimal depletion performance at substantially lower computation cost. EDF-WPT, the strongest reactive policy, incurs 2-4 times more depletion events than DT-PBB. Proactive DT-enabled look-ahead decisively outperforms reactive urgency-based scheduling, validating the zero-downtime paradigm for large-scale IIoT networks.

A. Alenezi · 0 citations
Open access Jul 2026

Reinforcement Learning for Resource Allocation in Energy-Harvesting Cooperative IoT Networks

This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.

Olumide Alamu, T. Olwal, Emmanuel M. Migabo · 0 citations