Skip to content
Book Open access

Experience with NVIDIA GPUDirect Storage (GDS) in Academic HPC Environments: Challenges, Pitfalls, and Practical Limitations

Jul 2026 · Practice and Experience in Advanced Research Computing · pp. 1-4 · 0 citations · 10 references
Computer Science

TL;DR

The results show that GDS performance depends heavily on file sizes, access patterns, and storage backends, and the path from benchmark to production is longer and more expensive than marketing materials suggest.

Abstract

The use of general-purpose GPUs has become essential in modern computing, particularly for workflows that are throughput-bound, e.g., deep learning workflows, LLMs, image segmentation, etc. Because GPUs function as coprocessors, data transfer between the CPU and GPU is constrained by the bandwidth of the interconnect. To mitigate this bottleneck, GPUDirect Storage (GDS) was introduced by Nvidia to enable more efficient data movement and improve overall system performance. GDS allows direct data transfers between storage and GPU memory, bypassing the CPU entirely. Vendors claim significant bandwidth improvements with minimal code changes. In this work, we report our experience deploying GDS on OSCAR, Brown University’s heterogeneous HPC cluster. We deployed and tested GDS across three storage configurations: VAST Data (pNFS over 200G HDR InfiniBand), IBM Spectrum Scale (GPFS over NDR InfiniBand), and local NVMe drives on DGX systems. We executed benchmarks both vendor-provided and a production workload using a Scientific Machine Learning (SciML) benchmark. Our results show that GDS performance depends heavily on file sizes, access patterns, and storage backends. The vendor benchmarks showed improvements in specific scenarios, but these gains did not always translate to the SciML benchmarks. Beyond the performance results, the deployment itself consumed months of staff time, multiple support tickets across NVIDIA, VAST, and IBM, and significant unplanned expenditure on dedicated optical components and cables that vendor planning documents never mentioned. The gains, where they existed, were modest. More critically, application-level support remains immature: PyTorch lists GDS integration as experimental, and at the time of writing, their own tutorial code has been removed from the documentation website. We discussed this issue with PyTorch developers via GitHub, but no clear solution was provided. The feature is currently labeled as experimental and does not appear to be under active development. By sharing our experience, we hope to give other HPC centers realistic expectations for GDS deployments. The technology works, but the path from benchmark to production is longer and more expensive than marketing materials suggest.

Read PDF

Similar papers

Book Open access Sep 2026

Analysis of shared memory between CPUs and GPUs

This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.

Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al. · 0 citations
Preprint Aug 2026

Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study

Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intranode) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a s...

E. Krishnasamy · 0 citations
Preprint Aug 2026

GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance

This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the underlying acceleration mechanism. The design accelerates PDSCH, PUSCH, SRS, PRACH, split-8 lower-PHY transforms, and O-RAN...

M. Pennybacker, Wanze Liu, A. Kharchenko et al. · 2 citations
Preprint Aug 2026

GPU implementation of a resource-constrained virtual machine

This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.

S. Li, Vladislav Brusokas, Andrei Ghita et al. · 0 citations

FitFloat: Read/Write Random-Access Compressed Floating-Point Arrays for GPUs

FitFloat is presented, a drop-in floating-point array replacement supporting user-specified precision on GPUs with the goal of reducing storage requirements of scientific applications while maximizing performance over Unified Memory.

Andrew Rodriguez, Martin Burtscher · 0 citations
Book Open access Aug 2026

A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU

This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.

Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.