There exist applications that benefit from this feature, making it attractive even for already ported applications, and its implementation in the software stack (OS kernel, compiler, and runtime) is introduced.
Abstract
OpenMP 5.0 introduced the Unified Shared Memory (USM) feature through the requires directive. The feature simplifies the adoption of the OpenMP programming model by providing a unique and common address space between the accelerators and the host and allowing the access (dereference) of the same memory address on different devices, thus avoiding the burden of explicit data transfers to maintain the consistency between the address spaces. Hence, the feature eases quick prototyping and porting of applications to OpenMP with accelerators. In this paper, we introduce the Intel implementation for USM. We briefly discuss its implementation in the software stack (OS kernel, compiler, and runtime), then assess its adoption complexity in existing HPC applications using OpenMP for accelerators, and, finally, evaluate the performance of these applications when adopting USM on an Intel Battlemage GPU. USM is not expected to grant performance uplifts to already optimized applications with explicit, granular data-motion control and our results show an overhead with a geometric mean below 1.2x (1.03x seems achievable with further optimizations). Yet, in this paper we show there exist applications that benefit from this feature, making it attractive even for already ported applications.
CUDA is the dominant GPU programming model in HPC and industrial accelerator software, and a large body of production code is written directly in it. Deploying that code on non-NVIDIA accelerators has traditionally required source translation, backend-specific rewrites, or a full rewrite in a new programming model. Thi...
Beau Johnston, Chris Kitching, Matthew Ireland et al.· Workshop Proceedings of the...· 0 citations
This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the underlying acceleration mechanism. The design accelerates PDSCH, PUSCH, SRS, PRACH, split-8 lower-PHY transforms, and O-RAN...
M. Pennybacker, Wanze Liu, A. Kharchenko et al.· 2 citations
It is found that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility.
R. Caplan, Miko M. Stulajter, J. Linker et al.· 0 citations
This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.
Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al.· Proceedings of the 4th Works...· 0 citations
(English) High-performance computing (HPC) platforms are evolving towards increasingly complex architectures: many-core CPUs with multi-level NUMA hierarchies, heterogeneity with multiple classes of accelerators and higher-capacity interconnects. The increasing complexity and variety of resources in these machines make...
This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.
S. Li, Vladislav Brusokas, Andrei Ghita et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.