XTLS: Scalable TLS Offloading through Host-SmartNIC Stack Co-Design
Abstract
Transport Layer Security (TLS) has become indispensable to modern network communication, yet its cryptographic overhead remains a well-known burden. Despite extensive acceleration of cryptographic operations, the TLS handshake continues to impose a substantial "tax" on host CPUs, consuming precious cycles that could otherwise be utilized for tenant workloads. This paper presents XTLS (eXtract TLS), a TLS offloading architecture that liberates the host CPU from this "TLS tax". We present a dual-stack design that offloads the entire TLS stack to the NIC—leaving the host to focus on application logic—and an adaptive per-flow symmetric-crypto NIC offload policy that jointly optimizes Requests Per Second (RPS) and data rates. We address the key design considerations on both sides: a scalable NIC-side handshake stack, robust inter-stack coordination between the SmartNIC and the host, and host-side portability for diverse L7 applications. As a result, XTLS delivers equivalent RPS performance using just a single host core, requiring 8× fewer host resources compared to the CPU-only approach and 7× fewer than Intel QAT. Despite this minimal resource usage, XTLS achieves 1.1× lower tail latency than QAT and 1.5× lower than CPU-only methods, all while sustaining the same data rates.