Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 15 references
Abstract
Dynamic wireless resource allocation in multi-cell networks is challenging due to non-stationary traffic, intercell interference coupling, and heterogeneous quality-of-service (QoS) constraints. Conventional schedulers and standalone metaheuristics lack adaptability across operating regimes, while deep reinforcement learning (DRL) methods often incur high training complexity and stability limitations. This paper proposes a context-aware reinforcement hyper-heuristic framework for dynamic wireless resource allocation. A contextual bandit controller hierarchically selects among multiple low-level optimization heuristics based on real-time network state features. A multi-objective reward design jointly optimizes throughput, fairness, power efficiency, and allocation stability. We establish sublinear regret guarantees under the contextual bandit model and prove convergence under standard stochastic approximation conditions. Extensive simulations over 5,000 large-scale multi-cell instances demonstrate consistent improvements over proportional fair scheduling, evolutionary methods, and DRL-based allocators in throughput, Jain's fairness index, convergence speed, and robustness to traffic perturbations. Statistical tests confirm the significance of the gains. The results indicate that reinforcementdriven hyper-heuristic orchestration provides a scalable and theoretically grounded solution for dynamic wireless resource management.
Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation. While multi-agent reinforcement learning (MARL) is a natural framework for such distributed sequential control, its application here faces two difficulties: finite-horizon budget constraints cannot be evaluated at each time slot, and the nonlinear proportional fairness utility admits no principled per-slot decomposition. We propose HeLyMARL, a Lyapunov-embedded heterogeneous MARL framework that resolves both via drift-plus-penalty decomposition with virtual queues. The energy and handover constraint pressures are internalized directly into a unified per-slot reward, converting the constrained finite-horizon problem into an unconstrained MARL problem. Comparison against two Lagrangian-based alternatives reveals a timescale separation: Lagrangian relaxation regulates constraints only across training episodes, whereas the virtual queues of HeLyMARL bound cumulative budget consumption at every partial horizon within an episode, a pacing guarantee beyond the reach of greedy Lyapunov-based control. Simulations show that HeLyMARL is the only method that sustains the throughput-fairness balance together with uninterrupted service throughout the horizon, outperforming conventional MARL, Lyapunov-based, and constrained MARL benchmarks without premature budget exhaustion.
Yeonseo Jeong, Wonhyeok Ko, Sungweon Hong et al.· 0 citations
The densification of wireless networks and growing real-time service demands have intensified the need for intelligent, energy-efficient resource allocation. Traditional static and centralized methods fall short in adapting to the dynamic and interference-prone nature of 5G and emerging 6G environments. This study proposes a decentralized reinforcement learning (RL)-based framework for joint power and spectrum allocation in ultra-dense wireless systems. Each base station acts as an autonomous agent, making real-time decisions based on local traffic and interference conditions. Simulated using a custom Python-based environment with 50 base stations and 500 users, the RL approach is benchmarked against static and optimization-based methods. Results show the RL model achieves up to 91% energy efficiency, 94% spectrum utilization, and only 5% QoS degradation, outperforming baseline models. This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.
Mugerwa Joseph, Ajaegbu Chigozirim· International Journal Of Eng...· 0 citations
The sixth generation (6G) of wireless networks must ensure high reliability, availability, and fairness, even in dynamic and challenging environments. Environmental-aware knowledge, such as that obtained from Radio Environment Maps (REMs), offers predictive insights into channel conditions and can guide more effective scheduling decisions. This paper proposes a scalable deep reinforcement learning (DRL) framework that exploits such knowledge to optimize multi-user scheduling under limited resources, in order to enhance reliability and availability while maintaining fairness. Unlike standard deep Q-network (DQN), which evaluates Q-values per action, we propose a novel learning method that estimates per-user Q-values, enabling user selection with per-decision complexity that scales linearly with the number of users. Simulation results show that the proposed approach consistently balances reliability and availability while maintaining fairness, outperforming Round Robin (RR) and Proportional Fair (PF) schedulers, especially in environments with high clutter density and frequent non-line-of-sight conditions. In particular, the proposed method improves reliability by over 400% with only a 2% drop in availability, compared to the RR scheduler. Relative to PF, it improves availability and fairness by 23% and 40%, respectively, without sacrificing reliability. These improvements are observed under balanced propagation conditions, where the probabilities of line-of-sight and non-line-of-sight are equal due to a clutter density of 50%. These results highlight the potential of the proposed environmental-aware DRL scheduler to support trustworthy 6G communication in complex and dynamic environments.
Roya Khanzadeh, Fjolla Ademaj-Berisha, Bernhard Etzlinger et al.· IEEE Transactions on Machine...· 0 citations
This formulation provides a rigorous and tractable framework for distributed spectrum sharing in 6G O-RAN systems, with the potential to support intelligent and adaptive control in future wireless networks.
E. Spyrou, Chrysostomos D. Stylios, V. Kappatos et al.· Future Internet· 0 citations
Managing dynamic User Association and Resource Allocation (UARA) in modern Heterogeneous Cellular Networks (HetNets) remains a critical open challenge. Existing mathematical optimization and Reinforcement Learning approaches face limitations in handling low-latency decision-making under dynamic traffic conditions. This paper introduces a novel orchestration scheme for game-theoretic UARA in HetNets. The proposed bilevel framework distributes UARA decisions to User Equipment through a multi-objective non-cooperative game. Overlaying the distributed game, a centralized Deep Reinforcement Learning controller orchestrates network performance by dynamically configuring the game's utility parameters, enabling transitions between power awareness, coverage enhancement, and balanced operation. Evaluated on urban HetNet topologies with 3GPP TR 38.901-compliant channel modeling, the proposed framework closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods. Furthermore, it incurs low computational overhead and maintains stable performance across the evaluated traffic densities without retraining.
Sotiris Kopsinos, Alexandros I. Papadopoulos, Antonios Lalas et al.· 0 citations