Skip to content
Book Open access

PLINK: A GPU-Initiated I/O Platform Exploiting NVMe Parallelism

Sep 2026 · Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems · pp. 71-76 · 0 citations · 4 references

TL;DR

PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence.

Abstract

As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps only a single command outstanding per thread; moreover, non-uniform per-thread I/O completion latency induces warp divergence that wastes a substantial fraction of GPU clock cycles. Guided by a cycle-granularity breakdown of BaM’s I/O path that localizes over 93% of I/O turnaround time to three phases, we present PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence. An evaluation on a state-of-the-art GPU paired with a PCIe Gen 6 NVMe SSD, in which each optimization is applied one at a time on top of the BaM baseline, shows that PLINK’s I/O backend reaches SSD IOPS saturation at 136 total warps—submission and completion combined—versus BaM’s 238.

Read PDF

Similar papers

Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations
Preprint Aug 2026

GPU Offload in Rust: Portable, Safe, and Fast

This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.

Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al. · 1 citation
Review Aug 2026

Concurrency Response of Plain Global Loads on the NVIDIA H100

The main result concerns the plain-load path: attained LDG bandwidth peaks at a small offered per-thread load (K ~ 2) and then declines, by about 35% from K=2 to K=8 at the authors' primary configuration.

S. Manjunath, Rahul Ramachandra · 0 citations
Book Open access Sep 2026

Anchor: Mitigating GPU Shallow Disruptions with Decoupled Memory

Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems,...

Hao-Yi Ma, Shi-Wei Gao, You-Min Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.