Aug 2026· Workshop Proceedings of the 55th International Conference on Parallel Processing· 0 citations· 12 references
Computer Science
TL;DR
This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.
Abstract
CPU-GPU coscheduling enables simultaneous execution of an application across both processing units, but its efficiency depends on workload partitioning and memory architecture. This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform. Using sparse conjugate gradient (CG) as a case study, we assess various work divisions across three memory-management paradigms: explicit copy, managed memory, and mapped memory. Our evaluation highlights the run time and programmability tradeoffs of reducing manual CPU–GPU data movement. The results show that compared with the H100 PCIe platform, GH200 makes several hybrid CPU–GPU work divisions competitive and makes managed memory practical for several matrices. These results suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.
This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.
Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al.· Proceedings of the 4th Works...· 0 citations
This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.
S. Li, Vladislav Brusokas, Andrei Ghita et al.· 0 citations
A three-tier buffer manager that manages encoded pages across SSD, RAM and GPU device memory while codec-specific operators fuse decompression, selection, join and aggregation is presented.
Maha Alwahibi, Maximilian E. Schüle· Datenbank-Spektrum· 0 citations
The advent of cloud-based artificial intelligence and the increased digitalization of embedded systems require powerful GPUs capable of simultaneously running kernels from different software providers. To accommodate the resource isolation and Execution Time Determinism (ETD) needed with the increasing number of kernel...
Vahid Geraeinejad, Paul Delestrac, Javier Barrera et al.· IEEE International Conferenc...· 0 citations
These findings demonstrate that high-level ECS simulation programs can be compiled into efficient GPU execution without requiring users to manually implement and coordinate low-level kernels.
This work explores the efficacy of the Chapel programming language’s GPU support for implementing irregular distributed GPU graph applications. Chapel’s partitioned global address space (PGAS) model provides a cohesive way to target distributed nodes, CPU concurrency and task parallelism, and both CPU and GPU single in...
Paul Sathre, Wu-Chun Feng· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.