It is found that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility.
Abstract
There continues to be growing interest in using standard language constructs for parallel and accelerated HPC computing, avoiding the need for (sometimes vendor-specific) external APIs. For Fortran applications, language features such as'do concurrent'loops open the door for compilers to implement multi-threaded, GPU-accelerated, and even distributed multi-node code with only the standard language. Here, we explore the current status of using'do concurrent'for GPU-accelerated Fortran applications across three major GPU vendors (NVIDIA, AMD, and Intel). Using a production application, we test their current capabilities, showing where the standard language alone can be used, and where augmenting the code with a directive-based API (e.g., OpenMP) is still desirable or required. Multi-GPU tests are performed with GPU-aware MPI libraries. We find that the three GPU vendors can now GPU-accelerate pure Fortran (zero directives), but that manual data movement directives can help with performance and compatibility. The results show that there is rapid advancement towards making GPU-accelerated scientific HPC code performance portable using the Fortran standard language.
This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.
S. Li, Vladislav Brusokas, Andrei Ghita et al.· 0 citations
Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intranode) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a s...
Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.
David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley· 0 citations
CUDA is the dominant GPU programming model in HPC and industrial accelerator software, and a large body of production code is written directly in it. Deploying that code on non-NVIDIA accelerators has traditionally required source translation, backend-specific rewrites, or a full rewrite in a new programming model. Thi...
Beau Johnston, Chris Kitching, Matthew Ireland et al.· Workshop Proceedings of the...· 0 citations
New AI accelerators arrive before the kernels that make them fast, because peak kernel performance requires architecture-specific expertise in operand pipelines and data-movement techniques. Coding agents can now write, compile, and tune kernels on their own, so they could greatly accelerate kernel development and opti...
Ye-Hong Jiang, Sheng Chen, Fang-Wen Fu et al.· 0 citations
We propose PKDB, the first interactive debugger for GPU and multithreaded low-level kernels written in Python. Python is widely used in high performance computing (HPC), with frameworks such as PyKokkos translating Python-embedded domain-specific languages to native code that runs across OpenMP-threaded CPUs and variou...
Ivan Grigorik, Gabriel Kosmacher, G. Biros et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.