Sep 2026· Proceedings of the International Conference on Parallel Processing· pp. 780-790· 0 citations· 10 references
Abstract
Serverless functions are short-lived, yet RDMA Reliable Connections (RCs) are expensive to create and assume long-lived endpoints. Even when compute sandboxes are warm, an invocation can still suffer a network-cold start if it establishes a new RC or accesses a remote endpoint for the first time. We present RSC (Reliable Serverless Connectivity), a high-performance user-space RDMA stack that eliminates RC setup from the invocation hot path. RSC exposes multiple operating modes that trade off latency, CPU efficiency, isolation, payload registration, and physical QP budget to support diverse serverless workloads. Its density-adaptive architecture scales efficiently to thousands of tenants, providing an in-process path for latency-critical calls, a pre-registered path for large transfers, memory-window isolation, and a diagnostic shared-agent path that measures the cost of physical QP consolidation. Our evaluation on 100 Gbps RoCE shows that native RC setup costs 7.94 ms, while reused-connection 64 B calls reach 59.87 μ s in adaptive sleep mode and 4.70 μ s with aggressive busy polling. For 2 MB payloads, the pre-registered path reaches 236.80 μ s. Workflow, request‑mix, resource, noisy‑neighbor, endpoint‑scale, and QP‑density diagnostics confirm that connection reuse removes millisecond‑scale RC setup from warm service invocation paths. These results demonstrate that RSC can make RDMA economically viable for serverless services.
This work presents ActiveRDMA, an active RDMA model that extends traditional one-sided semantics by leveraging on-NIC programmability, and enables complex, multi-step RDMA patterns to execute directly on the target side, eliminating host CPU involvement and reducing communication round-trips.
Jerónimo Sánchez García, Peter-Jan Gootzen, Raphael Frantz et al.· Proceedings of the Internati...· 0 citations
STORM is presented, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog, and converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued d...
Jichun Wu, Ran Shu, Gianni Antichi et al.· Conference on Applications,...· 0 citations
It is shown that TCP performance can be dramatically improved by following the design philosophy of RoCE, which integrates kernel bypassing, segment offloading, and zero-copy buffer management, and software TCP stacks can achieve performance parity with modern RNICs.
Yonghwan Chung, Yi-Han Dang, Kyoungsoo Park· Proceedings of the 17th ACM...· 0 citations
CERLA-SFC is introduced, a hierarchical, multi-objective orchestrator that unifies learning-based placement, topology-aware routing and event-driven resource allocation in a single control loop that maintains near-zero latency violations across all urgency classes while keeping end-to-end delay in the millisecond range...
Yuanfei Xiao, Zhenli He, Xiaolong Zhai et al.· 0 citations
This work introduces CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header, and proposes Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production.
Abhiram Ravi, Nandita Dukkipati, Weiwu Pang et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.