Skip to content

PPO-MS: Confidence-Aware and Collaborative Traffic Management for Multimedia Streaming in SDN

2026 · IEEE Transactions on Network and Service Management · Vol 23, pp. 6836-6848 · 0 citations · 41 references

Abstract

The rapid growth of multimedia streaming poses critical challenges, including bursty traffic and congestion, leading to playback delays. The existing separate prediction and control mechanisms for multimedia traffic scheduling, which are based on software-defined networks (SDN), are unable to proactively manage bursty traffic under uncertain conditions. This limitation is particularly evident in SDN-enabled backbone and multimedia-aware access networks, which typically assume centralized control and stable topologies. They lack integration of traffic prediction, traffic shaping, and real-time perception scheduling through reinforcement learning, resulting in low efficiency when exploring multiple paths in dynamic networks. To address this challenge, we propose PPO-MS (Proximal Policy Optimization-based Multimedia Scheduler), an SDN-based multimedia traffic scheduling algorithm integrating three key innovations: 1) A novel LSTM+HTB synergy where LSTM’s confidence intervals dynamically adjust HTB (Hierarchical Token Bucket) shaping parameters, enabling adaptive rate control under prediction uncertainty and overcoming the limitations of static LSTM+HTB hybrids; 2) A Deep Reinforcement Learning (DRL)-optimized path pruning method that reduces state and action spaces by generating a constrained set of $k$ disjoint candidate paths via an improved redundant tree algorithm. Unlike traditional multi-path schemes, this method tightly couples path preselection with the RL decision loop for adaptive, context-aware routing; 3) Generalized Advantage Estimation (GAE)-accelerated PPO for stable convergence in dynamic environments. In contrast to prior works (e.g., LSTM+RL for QoE or standalone tree algorithms), PPO-MS uniquely unifies these modules through confidence-aware traffic shaping and hierarchical decision-making, validated via comparative experiments. Results demonstrate that PPO-MS, through the synergistic integration of confidence-aware traffic shaping and DRL-optimized path pruning, significantly outperforms decoupled baselines. In particular, via isolation studies against simpler alternatives (e.g., mean-prediction and fixed-margin shaping), the confidence-aware shaping mechanism is validated to be superior under bursty traffic conditions. Overall, PPO-MS reduces end-to-end latency by 17.3% and packet loss by 32.4% while achieving 24.4% better load balancing during traffic bursts.

View source