Skip to content
Book Open access

Dorado: Scaling SmartNIC Session Tables on Commodity DDRs

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · 0 citations · 52 references
Computer Science

TL;DR

Dorado is a novel design that scales SmartNIC session tables entirely on inexpensive DDR modules and uses three new techniques that extract commodity DDR performance by restructuring session table layout, decomposing processing pipelines to reduce locking, and scheduling memory accesses to minimize stalls.

Abstract

FPGA-based SmartNICs are widely deployed for cloud network function acceleration, but their memory subsystem is under increasing pressure because of large session tables. Conventional wisdom suggests that high packet processing performance relies on advanced memories (e.g., SRAM, HBM), but those are costly to add at cloud scale. Dorado is a novel design that scales SmartNIC session tables entirely on inexpensive DDR modules. At the heart of Dorado are three new techniques that extract commodity DDR performance by restructuring session table layout, decomposing processing pipelines to reduce locking, and scheduling memory accesses to minimize stalls. Our testbed results show that Dorado improves packet processing rates by 33%, even with fewer hardware resources. Further, we have deployed Dorado to millions of servers, processing network traffic from billions of users on a large public cloud for over three years. Our production results show that Dorado can accommodate up to 16M session entries, reduce memory cost by 80%, while enabling 50Mpps line-rate processing.

Read PDF

Similar papers

Conference Aug 2026

MemMax: Memory-Parallel FPGA Optimization for Bandwidth-Bound IoT Image Processing

Memory-bound workloads increasingly dominate modern data-intensive systems, especially in Internet of Things (IoT) pipelines where large volumes of sensor and image data must be processed under strict latency and power constraints, yet CPUs quickly saturate their memory bandwidth even with many cores. FPGAs offer highe...

Benjamin Mikailenko, R. Rongon, Xiaokun Yang et al. · 0 citations

Finding NEMO: Nimble and Expressive Memory Observability

NEMO is presented, a nimble and expressive hardware memory telemetry engine for server memory controllers (MCs) that gives OS subsystems policy-specific views of memory behavior and provides higher-fidelity signals at substantially lower CPU overhead across a range of state-of-the-art memory management systems.

Shi-Hang Li, Matthew Giordano, Tushar Garg et al. · 1 citation
Preprint Aug 2026

Oasis: Hiding the Cost of Querying Parquet Files in the Datapath

Oasis is a data-processing SmartNIC that offloads Parquet decoding into the network datapath as a custom hardware accelerator, and shows that Oasis hides the cost of Parquet decoding behind the network datapath with minimal overhead, overlapping the scan with the remainder of the query execution.

Jonas Dann, Luca Tagliavini, Gustavo Alonso · 0 citations
Jul 2026

StrataCL: Fabric-Native Communication Library for Production Supernodes

StrataCL introduces registration-on-allocation to realize user-buffer direct communication, and designs communication operators with workload-balanced NPU-core partitioning and NPU-driven SDMA offloading to exploit supernode architecture features.

Tian-Cheng Hu, Jin Qin, Yu-Zheng Wang et al. · 0 citations
Aug 2026

Achieving High-Performance Erasure Code Repair Through DPU Offloading

Erasure coding provides efficient fault tolerance for large-scale distributed storage systems. However, its data repair process is well-known to be resource-intensive. We find that conventional host-centric, TCP-based repair architectures suffer from severe resource contention. Even in high-bandwidth networks, such int...

Xiangyu Yao, Yina Lv, Tianyu Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.