Jul 2026· Journal of Industrial Information Integration· Vol 53, pp. 101170· 0 citations· 44 references
Computer Science
TL;DR
This paper proposes the use of AQM as a lightweight and non-intrusive mechanism for assisting mission-critical traffic flows in IIoT networks and demonstrates that multi-queue AQM schemes provide substantial flow isolation and capacity sharing benefits, and significantly improve the performance of mission-critical traffic flows under network pressure.
Abstract
Mission-critical Industrial Internet of Things (IIoT) traffic flows require bounded network latency and jitter guarantees to ensure the safe functioning of critical industrial infrastructure. These flows are typically communicated via commodity network routers with conventional First-In-First-Out (FIFO) buffers. FIFO has proven to be the culprit of the well-known bufferbloat phenomenon, and the deployment of Active Queue Management (AQM) schemes have demonstrated significant performance improvements for latency-sensitive applications over the Internet in the IT domain. However, the bufferbloat phenomenon and the efficacy of AQM schemes have not been studied in IIoT-based OT domain. In this paper, we propose the use of AQM as a lightweight and non-intrusive mechanism for assisting mission-critical traffic flows in IIoT networks. Our experimental results demonstrated that multi-queue AQM schemes provide substantial flow isolation and capacity sharing benefits, and significantly improve the performance of mission-critical traffic flows under network pressure. We further provide deployment recommendations based on our experimental insights.
The rapid advancement of digital transformation (DX) requires networks to concurrently support diverse traffic types on a shared infrastructure. Next-generation infrastructures, such as the innovative optical and wireless network (IOWN) and high-speed data centers, require the coexistence of flows with disparate requirements. Here, low-delay critical communications (e.g., real-time control) must function alongside high-volume non-critical communications (e.g., background file transfers). While traffic shaping can effectively suppress the burstiness of critical flows, concurrent non-critical flows often exhibit significant temporal burstiness. In basic combined input and output queued (CIOQ) switches, such bursty flows can monopolize output FIFO buffers, causing significant delay degradation for critical flows. To mitigate this interference, this paper investigates a virtual input queue (VIQ) scheme, which provides logical isolation at the output stage. We evaluate its performance across diverse conditions via discrete-time simulations. Simulation results show that the CIOQ switch with VIQs effectively isolates critical flow from the impact of bursty flow originating from other input ports, thereby mitigating delay degradation caused by cross-port interference under the examined traffic conditions. They also reveal how the burstiness of non-critical flows, the number of switch ports, arrival rates, and traffic uniformity influence the performance.
Takuto Kubo, Shingo Okada, Eiji Oki· IEEE Open Journal of the Com...· 0 citations
As Data Center Networks (DCNs) continue to scale, the limitations of traditional centralized Software-Defined Networking (SDN) architectures become increasingly apparent, as they fail to meet the stringent demands for low latency and quality of service (QoS). In this paper, we propose an adaptive traffic-aware load balancing mechanism (ATL), a telemetrydriven in-switch scheme implemented on the programmable data plane (PDP) using P4 and driven by In-band Network Telemetry (INT). The current traffic regime is inferred by analyzing the remaining capacity (RC) of each link and its short-term variation (VAR), and adopts a dual-optimization strategy: (i) separating elephant flows (large flows) and mice flows (small flows) onto disjoint path sets to mitigate head-of-line blocking and packet reordering; (ii) dynamically adjusting the flowlet threshold $\left(F^{*}\right)$ to strike a balance between maximizing parallelism and ensuring in-order delivery. We prototyped and evaluated ATL in a Mininet/BMv2 environment, targeting bandwidth-constrained scenarios representative of IoT and edge deployments. The results show that, compared to existing methods such as ECMP, HULA, AWCMP, and APS, ATL consistently reduces both the average and 99th-percentile AFCT while achieving superior elephant-flow throughput, with notable improvements in traffic stability and packet-ordering preservation. Furthermore, ATL demonstrates a favorable cost-performance trade-off ratio of 1:0.99, confirming its efficiency and feasibility within the resource-constrained P4 switch environment.
Lossless fabrics are widely used in many production data centers, but they can give rise to issues such as head-of-line blocking and congestion spreading during network congestion, which can significantly degrade the performance of data center applications. Additionally, the latency of end-to-end solutions can lead to the buildup of switch queues. To address these challenges, this paper proposes a method called Priority Flow Control-Sensitive (PFC-S). PFC-S monitors buffer occupancy and traffic intensity, performs proactive rerouting, and prevents the impact of PFC congestion diffusion on victim flows. This approach helps maintain low buffer usage levels, thereby enabling control over tail latency. Initial evaluations demonstrate that PFC-S can reduce the average flow completion time and effectively prevent congestion spreading. Moreover, experimental results show that PFC-S provides better protection for victim flows compared to standard PFC, BFC, and HPCC methods.
Weimin Gao, Jiawei Huang, Qile Wang et al.· Journal of High Speed Networ...· 0 citations
In data centers, large-scale many-to-one traffic can rapidly exhaust switch buffers and trigger priority-based flow control (PFC) pause, resulting in increased flow completion time (FCT) for uncongested flows. To address this issue, we propose an innovative switch-side fast and accurate flow control (FAFC) scheme. By differentially allocating pause time for each port during congestion, FAFC can minimize the performance loss for uncongested flows. Furthermore, FAFC is also coupled with an effective queue length prediction algorithm to enable proactive and reliable estimation of the congestion level. Extensive system-level simulations demonstrate that FAFC can flexibly allocate pause times across congested ports, which are not only compatible with existing PFC but also do not require per-flow states. We implemented FAFC in P4 programmable switches, showing it as lightweight flow control method that is portable for implementation in hardware. Remarkably, our large-scale simulations illustrate that compared to traditional PFC, FAFC improves the average FCT slowdown and 95% FCT slowdown by 10.6% and 23.3%, respectively, under Hadoop workload when performing HPCC congestion control.
Chengdi Lu, Yuang Chen, Fangyu Zhang et al.· IEEE Transactions on Network...· 0 citations
In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flows. However, before throughput degradation becomes observable, the network often exhibits early signs characterized by rising flow activity and queue overflow signals. Existing congestion-management approaches primarily react only after congestion becomes visible, leaving these early signs largely unexploited. In this paper, we propose ProFlow, a proactive flow-placement framework for protecting performance-sensitive traffic in multi-tenant datacenter networks, thereby utilizing the early signs of potential throughput degradations. ProFlow leverages distributed telemetry signals and offline-trained reinforcement learning (RL) to identify precursor congestion conditions and proactively reroute protected flows before throughput degradation occurs. Evaluation results using FABRIC testbed show that ProFlow achieves approximately 40% higher mean throughput than a reactive rerouting baseline while initiating rerouting decisions around 34 seconds earlier on average, demonstrating the effectiveness of anticipatory congestion management.
Sourya Saha, Md. Nurul Absur, S. Debroy· 0 citations
3GPP Release 16 enables a 5G system to operate as a transparent IEEE 802.1 TSN bridge, but its scalability under heterogeneous industrial workloads remains insufficiently characterised. This paper uses the nascTime framework on OMNeT++/Simu5G to evaluate how many TSN endpoints a single 5G NR cell can bridge before per-flow QoS degrades. We model closed-loop control, machine vision, bulk telemetry, and IEEE 802.1AS traffic over a four-bearer SDAP architec- ture, varying the number of endpoints from 1 to 40, MAC scheduler, radio bandwidth (10 MHz and 20 MHz), and channel model. Results show three operating regimes. Below saturation, non-DRR schedulers perform similarly; near saturation, QoS- aware PF reduces critical-flow P99 latency by up to two or- ders of magnitude relative to channel-aware and fairness-based schedulers; and under overload, only QoS-PF maintains near- complete delivery for the highest-priority traffic. Across the two evaluated bandwidths, the saturation threshold approximately doubles when bandwidth doubles. We also show that isolating IEEE 802.1AS/gPTP traffic on a dedicated high-priority radio bearer reduces clock-servo instability, although endpoints carry- ing lower-priority data still experience elevated synchronisation delay under saturation because of reduced MAC scheduling frequency. Finally, the evaluated sub-6 GHz, 30 kHz-SCS con- figuration exhibits an effective latency floor of approximately 2.25 ms, indicating that sub-3 ms TSN deadlines may require radio-configuration changes such as configured grants or higher numerology
Mohamed A. M. Seliem, U. Roedig, C. Sreenan et al.· 0 citations