Machine Learning-Based Early Warning System for Transport-Layer Bottlenecks in Open-Source 5G Testbeds
Abstract
The transition toward disaggregated, software-defined 5G Standalone (SA) infrastructures has shifted much of the end-to-end latency budget from the radio interface into the transport and computational layers, where bufferbloat and congestion-window collapse degrade Quality of Service before any packet is lost. Contemporary congestion control remains fundamentally reactive, acting only after a bottleneck has materialized. This study designs and evaluates a machine learning-based early warning system that anticipates transport-layer bottlenecks several seconds in advance. We built an isolated, reproducible 5G SA testbed using Open5GS and srsRAN, emulating the air interface through ZeroMQ so that every variation in round-trip time, jitter, and throughput is attributable to queuing and protocol dynamics rather than radio-frequency noise. A stochastic generator injected variable multi-user loads over runs of up to six hours, yielding open one-second telemetry. Round-trip-time forecasting was framed as multivariate, multi-horizon regression under a strict honest-forecasting protocol that prevents temporal leakage. Across Ridge, Random Forest, HistGradientBoosting, GRU, and LSTM models, tree ensembles forecast short horizons accurately (R2≈0.90 at one second), temporal memory improves mid-range horizons, and all models converge to the trivial baseline at twenty seconds. As a binary early-warning task, the system catches most impending bottlenecks at high precision at five seconds. The dataset and pipeline are fully documented to support reproducibility and provide a transparent benchmark for proactive, zero-touch network orchestration.