Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 53 references
TL;DR
This work presents ActiveRDMA, an active RDMA model that extends traditional one-sided semantics by leveraging on-NIC programmability, and enables complex, multi-step RDMA patterns to execute directly on the target side, eliminating host CPU involvement and reducing communication round-trips.
Abstract
Modern distributed systems rely on one-sided RDMA for low-latency, high-throughput data movement without CPU intervention. However, one-sided RDMA provides only basic primitives (read, write, atomics) where the target side is passive. Applications with complex communication patterns, such as remote data structure traversal, require the host CPU to orchestrate multiple steps of communication, resulting in several network round-trips and thus diminishing the benefits of one-sided RDMA. We present ActiveRDMA, an active RDMA model that extends traditional one-sided semantics by leveraging on-NIC programmability. This enables complex, multi-step RDMA patterns to execute directly on the target side, eliminating host CPU involvement and reducing communication round-trips. We implement and evaluate ActiveRDMA on the NVIDIA DPA, an on-path SmartNIC architecture. Our evaluation demonstrates substantial improvements for fine-grained, latency-sensitive communication patterns: completion time reduces by up to 43% for messages up to 4 KiB that cannot benefit from batching. For operation rate, ActiveRDMA trades single-QP efficiency for scale-out throughput, surpassing conventional one-sided RDMA from 4+ connections onward, achieving up to 1.8× higher operation rate as the number of QPs scales. Bandwidth-dominated communications, however, see no advantage over conventional one-sided RDMA. In a distributed graph traversal application, we achieve up to 33% runtime reduction, with complete overlap of communication and computation.
Results demonstrate that RSC can make RDMA economically viable for serverless services, and workflow, request‑mix, resource, noisy‑neighbor, endpoint‑scale, and QP‑density diagnostics confirm that connection reuse removes millisecond‑scale RC setup from warm service invocation paths.
Wen-Di Song, Guang-Ping Xu, Yang-Yang Fan et al.· Proceedings of the Internati...· 0 citations
LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data...
PReCCL is a drop-in NCCL replacement that combines software inband telemetry with cross-VT workload reallocation, and implements in-band monitoring within the CCL, and precisely measures the stall counts of each VT, and piggybacks the telemetry meta-data on existing collective traffic.
Zhiyong Chen, Kaihui Gao, Li Chen et al.· Conference on Applications,...· 0 citations
STORM is presented, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog, and converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued d...
Jichun Wu, Ran Shu, Gianni Antichi et al.· Conference on Applications,...· 0 citations
Transport Layer Security (TLS) has become indispensable to modern network communication, yet its cryptographic overhead remains a well-known burden. Despite extensive acceleration of cryptographic operations, the TLS handshake continues to impose a substantial "tax" on host CPUs, consuming precious cycles that could ot...
Seongjong Bae, Junghan Yoon, Kyoungsoo Park· Proceedings of the 17th ACM...· 0 citations
It is shown that TCP performance can be dramatically improved by following the design philosophy of RoCE, which integrates kernel bypassing, segment offloading, and zero-copy buffer management, and software TCP stacks can achieve performance parity with modern RNICs.
Yonghwan Chung, Yi-Han Dang, Kyoungsoo Park· Proceedings of the 17th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.