QL-6GRP (Q-Learning for 6G Routing Protocol), a lightweight Q-learning-based routing protocol designed for fully distributed MANET environments, is presented, revealing key trade-offs between feedback frequency, routing quality, and computational scalability.
Abstract
Mobile Ad Hoc Networks (MANETs) are expected to support highly dynamic and decentralized communication scenarios in future 6G-oriented wireless systems. However, routing remains challenging because of mobility, topology variability, and resource constraints. Reinforcement learning (RL) offers a promising alternative by enabling adaptive routing decisions based on observed network conditions. This paper presents QL-6GRP (Q-Learning for 6G Routing Protocol), a lightweight Q-learning-based routing protocol designed for fully distributed MANET environments. The protocol enables each node to learn next-hop forwarding decisions using local observations, including link quality, residual energy, hop progress, and neighborhood density. A complete implementation of QL-6GRP was developed within the NS-3 simulator, supporting online learning through hop-level feedback signaling and bounded-memory operation. The protocol was evaluated under multiple parameter settings and network sizes using Random Waypoint mobility and UDP constant-bit-rate traffic to examine both routing performance and computational behavior. The experimental results demonstrate the feasibility of adaptive routing with moderate signaling overhead under carefully tuned moderate-scale scenarios while revealing key trade-offs between feedback frequency, routing quality, and computational scalability. Moderate periodic feedback provides the most favorable balance, whereas excessive feedback increases overhead without improving performance. In addition, reinforcement-learning operations incur substantial computational costs, with the wall-clock runtime increasing by approximately 13 times when the network size increases from 50 to 100 nodes. These findings reveal the operating limits and practical design trade-offs of lightweight tabular RL-based MANET routing and provide useful guidelines for future scalable learning-driven protocols in dynamic wireless environments.
A two-level Q-learning-based geographic routing protocol called TLQ-Geo for FANETs, which significantly reduces convergence time and computational overhead and integrates hierarchical decision-making with adaptive reinforcement learning.
Mehdi Hosseinzadeh, Jawad Tanveer, Amir Masoud Rahmani et al.· Journal of King Saud Univers...· 0 citations
An Adaptive Hybrid Routing Framework that integrates RPL and GPSR under a machine learning (ML)-driven decision engine that offers a resilient and energy-efficient routing solution for next-generation smart grid neighborhood area networks is proposed.
Teslim Komolafe, Enoch Owoeye, Samuel A. Adegbola et al.· Energy Science, Engineering,...· 0 citations
– Mobile Ad Hoc Networks (MANETs) play a critical role in disaster recovery, military communications, vehicular networking, emergency response, and remote monitoring applications. However, sustaining Quality of Service (QoS) in MANETs is challenging because of changing topologies, mobile nodes, energy constraints, unreliable wireless links, and congestion. Widely adopted routing protocols, including Ad hoc On-Demand Distance Vector (AODV), Dynamic Source Routing (DSR), Destination-Sequenced Distance Vector (DSDV), and Optimized Link State Routing (OLSR), rely on reactive or proactive route discovery mechanisms and often fail to adapt efficiently to rapidly changing network conditions. Although Deep Reinforcement Learning (DRL)-based routing approaches have demonstrated improved adaptability, they suffer from convergence instability, high computational complexity, continuous online learning requirements, and limited interpretability, restricting their practical deployment in mission-critical environments. To overcome these challenges, this paper presents an Explainable Neural Network-Driven Context-Aware Predictive Routing (XNN-CPR) method. The proposed method predicts the reliability of network links, selects routes based on QoS requirements, and uses SHAP to explain routing decisions. All these functions are combined into a single routing framework. The proposed method uses contextual network parameters, which include node mobility, residual energy, queue occupancy, Received Signal Strength Indicator (RSSI), and Signal-to-Noise Ratio (SNR), to infer future link reliability and congestion conditions. An explainable neural network is trained using simulation-generated data to estimate route stability, while SHAP-based explanations are employed to interpret and validate routing decisions. Unlike conventional DRL-based routing approaches that require continuous policy optimization during network operation, the proposed framework separates model training from routing execution, thereby reducing computational overhead and improving decision stability. Simulation experiments conducted in the NS-3 environment show that the proposed XNN-CPR framework steadily outperforms AODV, DSR, OLSR, and Deep Deterministic Policy Gradient (DDPG)- based routing protocols. Compared with the DRL-based benchmark, XNN-CPR obtains a 3.94% improvement in Packet Delivery Ratio (PDR), a 17.67% reduction in End-to-End Delay, and an 8.69% rise in Throughput. The results make sure that the integration of predictive intelligence with explainable decision-making noticeably enhances routing reliability, network stability, and QoS performance while ensuring transparency and trustworthiness in highly dynamic MANET environments.
Awadhesh Kumar Rai, Akhilesh A. Waoo· International Journal of Com...· 0 citations
The paper introduces RML-ZEREM to solve existing limitations, which functions as a Reinforcement Learning (RL) based Zone-Based Leader-Aware Energy-Efficient Routing Protocol for MANETs, which serves next-generation MANET applications.
Rani Sahu, Babita Rathore· Journal of Intelligent Compu...· 0 citations
The experiments show that feasibility-aware learning can approach deterministic baseline reliability while retaining learned forwarding capability under hop constraints, and confirm that action masking is the dominant mechanism for maintaining feasible routing decisions, whereas trust mainly provides reliability-aware regularization.
Adeel Iqbal, Muhammad Faisal Siddiqui· Computers, Materials & C...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.