Skip to content
Review Open access

Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review

Jun 2026 · Future Internet · Vol 18, pp. 340 · 0 citations · 145 references

TL;DR

According to this study, PPO provides continuous action spaces with good training stability for AI models and its stable policy-learning capabilities make it suitable for next-generation communication systems.

Abstract

Fifth-generation (5G), Beyond 5G (B5G), and sixth-generation (6G) wireless networks, along with the Internet of Things (IoT), are core communication infrastructure in smart cities. Their increased deployments create high-dimensional optimization and resource management challenges. Consequently, researchers have increasingly explored the use of Artificial Intelligence (AI) models for optimizing networks. The Proximal Policy Optimization (PPO) is one such algorithm that optimizes networks. This Systematic Literature Review (SLR) follows the PRISMA 2020 protocol to review 76 studies published between 2023 and 2026 to synthesize recent PPO-based approaches to optimize communication systems. This study examines key PPO variants in major communication domains. It outlines the primary obstacles to real-world deployment and provides a cross-domain classification. According to this study, PPO provides continuous action spaces with good training stability for AI models. Its stable policy-learning capabilities make it suitable for next-generation communication systems. However, sim-to-real transfer, reward design, and multi-agent scalability are a few key challenges encountered. Future directions emphasize robust, deployable PPO frameworks for 6G, IoT, and internet architecture.

Read PDF

Similar papers

Review Open access Jul 2026

Towards Intelligent 6G Networks: A Comprehensive Review of AI-Driven Control and Optimization

The transition from fifth-generation (5G) to sixth-generation (6G) communication networks represents a fundamental shift from conventional model-driven architectures toward AI-native, intelligence-driven network ecosystems capable of autonomous control, optimization, and decision-making. Although artificial intelligence (AI) has demonstrated significant potential in enhancing network performance, existing research remains fragmented, with most studies focusing on isolated network functions rather than integrated system-level intelligence. This study presents a systematic literature review (SLR) combined with a critical synthesis to examine the current state of AI-driven control and optimization in 6G communication networks. The review systematically analyzes 87 peer-reviewed studies published between 2020 and 2024, retrieved from IEEE Xplore, ScienceDirect, SpringerLink, and Wiley Online Library using predefined inclusion and exclusion criteria. The selected studies are critically evaluated with respect to machine learning, deep learning, and reinforcement learning techniques, emphasizing their architectural roles, operational capabilities, deployment feasibility, and system-level implications. The findings indicate that while AI techniques substantially improve network adaptability, resource management, and autonomous operation, significant challenges remain regarding scalability, computational complexity, data dependency, interoperability, explainability, and deployment in real-world environments. Furthermore, the review identifies a considerable gap between algorithmic advances and practical implementation, highlighting the need for integrated AI frameworks and architecture-aware design strategies capable of supporting scalable, trustworthy, and autonomous 6G communication systems. By providing a comprehensive synthesis of recent research, comparative analysis of major AI paradigms, and future research directions, this review contributes to bridging the gap between theoretical developments and practical deployment of AI-native communication networks.

Ali Ahmed Mirza, T. A. Mahmood, E. Dhulkefl · 0 citations
Conference Jul 2026

Learning-Based Resource Allocation in 5G NR Mode-2 Sidelink for Industrial AGV and AMR Communications

As the manufacturing sector increasingly adopts Industry 4.0 technologies, the need for reliable communication among devices, such as autonomous mobile robots (AMRs), or automated guided vehicles (AGVs), becomes a fundamental necessity. To support direct device-to-device communication, the third-generation partnership project (3GPP) introduced 5G new radio sidelink communication mode 2 (NR-SL) in Release 16, and 17. NR-SL allows devices to select transmission resources based on local channel sensing. In NR-SL, however, autonomous resource selection can lead to collisions, particularly in dense industrial environments. In this paper, we propose a learningbased resource allocation scheme (LBRA) for 5G NR Mode-2 sidelink. LBRA is designed to support industrial AGV and AMR communications. It employs multi-agent reinforcement learning (MARL) to improve resource allocation in NR-SL and reduces the collision probability. The results show that LBRA decreases the collision probability by approximately 73% compared to NR-SL.

Mahmoud Elsharief, Kiwoong Park, Han-Shin Jo · 0 citations
Review Open access Jul 2026

Evolution of Power Allocation Techniques in NOMA: Advancing 5G Toward 6G Networks

The transition of 5G and beyond wireless networks toward intelligence-driven and autonomous operation has revitalized strong interest in Non-Orthogonal Multiple Access (NOMA) as an efficient multiple access framework. Power allocation critically governs NOMA performance, directly impacting throughput, user fairness, and SIC effectiveness. This survey presents a focused review of power allocation strategies in NOMA, with emphasis on the progression from static and optimization-based dynamic schemes to data-driven Artificial Intelligence (AI) and Machine Learning (ML) driven approaches. In contrast to conventional strategies that require instantaneous channel state information and iterative optimization, AI/ML techniques enable adaptive, scalable, and low-latency decision-making in highly dynamic and nonconvex environments. Recent advances in reinforcement learning and deep learning for NOMA power control are discussed, highlighting key challenges like imperfect CSI, inter-cluster interference, and distributed learning constraints. This survey provides a concise AI-centric analysis and identifies promising directions for a practical learning-driven NOMA power allocation framework for future wireless networks. A consolidated, critically comparative analysis of NOMA power allocation that bridges the gap between 5G practice and 6G imperatives is also presented in this survey.

Lekshmi Nair M, Neelakantan Pc · 0 citations
Conference Jul 2026

LLM-Assisted Network Management for Digital Twin-Enabled 5G Systems

Efficient long-term network evolution is becoming increasingly critical in dense 5G-Advanced and beyond cellular systems, where persistent traffic imbalances and localized congestion pose significant challenges that conventional short-term radio resource management alone cannot fully mitigate. This paper proposes a digital twin (DT)-enabled non-real-time (NRT) network evolution framework integrated with a large language model (LLM). Within this architecture, the digital twin provides a high-fidelity, controllable environment for evaluating infrastructure actions, while the LLM serves as a strategic orchestration engine that recommends cost-efficient network upgrades based on observed network states. Unlike traditional optimization methods that require exhaustive mathematical reformulations for each specific scenario, the proposed framework leverages the reasoning capabilities of LLMs to interpret operator objectives and constraints in natural language, generating structured evolution plans. The considered NRT action space encompasses antenna upgrades, bandwidth expansion, and new base station (BS) deployment. A techno-economic formulation is introduced to jointly evaluate load reduction performance and overall economic expenditure. Numerical results in a dense cellular scenario demonstrate that the framework effectively reduces peak resource utilization and provides diverse, coordinated evolution strategies tailored to varying network conditions.

Yukai Wang, Janghee Woo, G. Hahm et al. · 0 citations