Skip to content

Automatic Model Compression and Quantized Deployment of Convolutional Neural Networks on Programmable Data Planes

2026 · IEEE Transactions on Networking · Vol 34, pp. 6765-6779 · 0 citations · 44 references

Abstract

The rapid development of programmable network devices and the widespread adoption of machine learning (ML) in networking have facilitated efficient research into intelligent data planes (IDPs). Offloading ML to programmable data planes (PDPs) enables quick analysis and responses to network traffic dynamics, and efficient management of network links. Compared to using an external low-cost board with sufficient memory and a general-purpose CPU, IDP deployment keeps inference inside the switch forwarding pipeline, avoiding inter-device transfer and coordination overhead. This enables line-rate processing and faster response for real-time network control. However, the hardware pipeline presents significant resource limitations. For instance, Intel Tofino ASIC has only 10Mb SRAM in each stage, and lacks support for multiplication, division, and floating-point operations. These constraints significantly hinder the development of IDP. This paper presents Quark, a framework that automatically compresses the convolutional neural network (CNN) and fully offloads quantized inference onto PDP. Quark employs model pruning to simplify the CNN model, uses quantization to support floating-point operations, and utilizes neural architecture search to balance accuracy and PDP resource constraints. Additionally, Quark divides the CNN into smaller units to improve resource utilization on the PDP. We have implemented a testbed prototype of Quark on both P4 hardware switch (Intel Tofino ASIC) and software switch (i.e., BMv2). Extensive evaluation results on the ISCX Botnet dataset demonstrate that Quark achieves 97.3% accuracy while using only 24.27% of the SRAM resources on the Intel Tofino ASIC switch, completing inference tasks at line rate with an average latency of $42.66\mu s$ .

View source

Similar papers

Conference Open access 2026

Key Algorithms of Convolutional Neural Networks and Hardware Implementation of Image Processing

Low-level hardware acceleration strategies to deconstruct the mapping from algorithm logic to silicon substrates are reviewed to provide strong guidelines for the hardware-software co-design of emerging ultra-low power edge Artificial Intelligence (AI) chips.

Linenxu Zhang · 0 citations
Open access 2023

Accelerating Neural Networks with Model Compression Techniques

Experimental results demonstrate that effective compression significantly reduces model size and computational cost with minimal performance loss, highlighting the importance of compression-aware design and concluding as a valuable reference for building efficient and scalable AI systems.

Daniel Rodríguez · 0 citations
Open access 2026

A High-Speed, Low-Power Pipelined CNN Implementation on FPGA Using On-Chip Processing and Layer-Wise Streaming Control

A high-speed CNN classification model with full layer-wise control is proposed, enabling seamless streaming of pixel data across all CNN layers via line buffers, resulting in high classification speed, low-power, and low energy consumption per classification.

Jyoti Pandey, Abhijit R. Asati · 0 citations
Aug 2026

Hardware-Aware Neural Network Deployment on Multi-Core in-Memory Computing Systems: A Compiler Perspective

Conventional compute-in-memory (CIM) deployment flows usually assume a fixed crossbar geometry, although DNN layers often have different channel counts and matrix shapes. The resulting shape mismatch leaves part of the array capacity unused and can increase the number of split-and-transfer operations during inference....

Kaiwen Deng, Sifan Sun, Hanjie Liu et al. · 0 citations
2026

An FPGA Integrated Programmable Switch Architecture for Data Plane DNN Inference

Machine learning (ML) is increasingly used in network data planes for advanced traffic analysis, but existing solutions (such as FlowLens, N3IC, BoS) still struggle to simultaneously achieve low latency, high throughput, and high accuracy. To address these challenges, we present <inline-formula> <tex-math notation="LaT...

Xiang-Yu Gao, Tong Li, Yin-Chao Zhang et al. · 0 citations
Aug 2026

A Flexible Framework for Layer-Parallel CNN Training on FPGA Clusters

A flexible and scalable hardware/software framework for training CNNs on Ethernet-connected FPGA clusters using tightly pipelined layer parallelism, highlighting the potential of FPGA-based, layer-parallel training for separable-convolution-dominated CNNs.

Philipp Kreowsky, Justin Knapheide, B. Stabernack · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.