Sep 2026· Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems· pp. 71-76· 0 citations· 4 references
TL;DR
PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence.
Abstract
As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps only a single command outstanding per thread; moreover, non-uniform per-thread I/O completion latency induces warp divergence that wastes a substantial fraction of GPU clock cycles. Guided by a cycle-granularity breakdown of BaM’s I/O path that localizes over 93% of I/O turnaround time to three phases, we present PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence. An evaluation on a state-of-the-art GPU paired with a PCIe Gen 6 NVMe SSD, in which each optimization is applied one at a time on top of the BaM baseline, shows that PLINK’s I/O backend reaches SSD IOPS saturation at 136 total warps—submission and completion combined—versus BaM’s 238.
Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.
David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley· 0 citations
This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.
Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al.· 1 citation
The main result concerns the plain-load path: attained LDG bandwidth peaks at a small offered per-thread load (K ~ 2) and then declines, by about 35% from K=2 to K=8 at the authors' primary configuration.
InplaceKVCache is proposed, the first KVCache abstraction whose format fixes each byte's physical residency at write time, so that the CPU--GPU load balance can be adjusted without moving data after placement, turning load balancing into pure scheduling.
Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems,...
Hao-Yi Ma, Shi-Wei Gao, You-Min Chen et al.· Proceedings of the ACM SIGOP...· 0 citations