Skip to content

An FPGA Integrated Programmable Switch Architecture for Data Plane DNN Inference

2026 · IEEE Transactions on Networking · Vol 34, pp. 6859-6874 · 0 citations · 73 references

Abstract

Machine learning (ML) is increasingly used in network data planes for advanced traffic analysis, but existing solutions (such as FlowLens, N3IC, BoS) still struggle to simultaneously achieve low latency, high throughput, and high accuracy. To address these challenges, we present <inline-formula> <tex-math notation="LaTeX">$\textsf {FENIX}$ </tex-math></inline-formula>, a hybrid in-network ML system that performs feature extraction on programmable switch ASICs and deep neural network inference on FPGAs. <inline-formula> <tex-math notation="LaTeX">$\textsf {FENIX}$ </tex-math></inline-formula> introduces a Data Engine that leverages a probabilistic token bucket algorithm to control the sending rate of feature streams, effectively addressing the throughput gap between programmable switch ASICs and FPGAs. In addition, <inline-formula> <tex-math notation="LaTeX">$\textsf {FENIX}$ </tex-math></inline-formula> designs a Model Engine to enable high-accuracy deep neural network inference in the network, overcoming the difficulty of deploying complex models on resource-constrained switch chips. We implement <inline-formula> <tex-math notation="LaTeX">$\textsf {FENIX}$ </tex-math></inline-formula> on a programmable switch platform that integrates a Tofino ASIC and a ZU19EG FPGA directly, and evaluate it on real-world network traffic datasets. Our results show that <inline-formula> <tex-math notation="LaTeX">$\textsf {FENIX}$ </tex-math></inline-formula> achieves microsecond-level inference latency and multi-terabit throughput with low hardware overhead, and delivers over 90% accuracy on mainstream network traffic classification tasks, outperforming the state of the art.

View source

Similar papers

2026

Automatic Model Compression and Quantized Deployment of Convolutional Neural Networks on Programmable Data Planes

The rapid development of programmable network devices and the widespread adoption of machine learning (ML) in networking have facilitated efficient research into intelligent data planes (IDPs). Offloading ML to programmable data planes (PDPs) enables quick analysis and responses to network traffic dynamics, and efficie...

Xiaoquan Zhang, Mai Zhang, L. Cui et al. · 0 citations
Open access Aug 2026

A Compiler-Aware Framework for Partitioned Neural Network Inference on FPGA DPUs

This work analyzes the Vitis AI compiler and proposes an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping, and restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection.

Federico Buccellato, Luca Mannini, C. De Sio · 0 citations
Preprint Sep 2026

Integer Quantization of Graph Neural Networks for Real-Time FPGA Track Finding

Real-time track finding for displaced-muon signatures in the CMS Level-1 trigger must operate under strict fixed-latency constraints of 12.5 $\mu$s while processing high-throughput detector data. Because muon hits map naturally onto sparse, irregular graphs, graph neural networks (GNNs) are attractive candidates; howev...

A. Cardini, P. Leguina, E. Aller et al. · 1 citation
Review Open access Aug 2026

Analysis of Research Progress on Deployment Methods for Deep Learning Models on FPGAs

A systematic review of FPGA-based DL deployment from a cross-layer perspective spanning model, compiler, architecture, runtime, and electronic design automation (EDA) is presented, highlighting that reliable cross-study comparison requires careful consideration of model configuration, precision, execution phase, batch...

Shuo Wang, Lei Chen, Chunsheng Tian et al. · 0 citations
Book Open access Sep 2026

PipeTree: Decision Tree Partitioning for Efficient Inference on Programmable Switches

In recent years, there has been a growing trend in deploying machine learning models directly in the data plane, taking advantage of the high throughput and low latency offered by modern programmable switches to conduct various tasks such as line-rate traffic classification and anomaly detection. Among these models, de...

Jia-Wei Huang, Xin Li, Yi-Jun Li et al. · 0 citations
Conference Aug 2026

Designing and Building an FPGA Accelerator That Uses Less Energy for DNN Inference

Deep Neural Networks (DNNs) are critical to modern AI applications, yet their deployment on standard CPUs and GPUs is constrained by high power consumption and computational latency, particularly in resource-constrained edge environments. To address these limitations, this paper presents the design and implementation o...

P. V. G. K. Rao, Dudekula Raziya · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.