Skip to content
Open access

Comparative Performance Analysis of Workload on Enterprise GPUs with Consumer Platforms Accelerated by CUDA Graphs

Jul 2026 · Anais do LIII Seminário Integrado de Software e Hardware (SEMISH 2026) · pp. 908-913 · 0 citations · 9 references

Abstract

This work investigates the feasibility of reproducing benchmarks originally run on datacenter GPUs such as the NVIDIA A100 and RTX 8000 using consumer-grade graphics cards, focusing on the NVIDIA GeForce RTX 3050 and GTX 1060 with CUDA Graphs support. Seven NAS Parallel Benchmarks (BT, LU, SP, EP, IS, MG, and CG) are evaluated across problem classes W, A, B, and C. Results show that the RTX 3050 delivers stable performance, typically 6×–12× slower than the A100, while VRAM limitations severely constrain the GTX 1060 for larger-scale problems. Although enterprise GPUs remain essential for massive, memory-bound workloads, modern consumer hardware combined with CUDA Graphs enables economical reproduction of moderate scientific experiments, supporting the democratization of high-performance computing research.

Read PDF

Similar papers

Book Open access Sep 2026

Analysis of shared memory between CPUs and GPUs

This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.

Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al. · 0 citations
Conference Open access Sep 2026

Analyzing the Impact of Architectural Design Decisions on Performance Across Generations of NVIDIA GPUs

This project analyzes how the architectural changes introduced across NVIDIA’s Volta, Ampere, and Hopper GPU generations translate into real-world performance and energy efficiency gains and shows that specialization can deliver substantial performance and efficiency gains, but only when workloads can effectively use t...

Matthew Tindale, I. Karlin, Tobias Salamon et al. · 0 citations
Book Open access Aug 2026

A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU

This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.

Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah et al. · 0 citations
Preprint Aug 2026

GPU implementation of a resource-constrained virtual machine

This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.

S. Li, Vladislav Brusokas, Andrei Ghita et al. · 0 citations
Preprint Aug 2026

Over the Memory Wall, Into the Instruction Wall: The New Bottleneck in GPU Data Processing

Valk, a performance analysis tool that combines data from multiple profilers, shows that when memory bandwidth is increased, kernels become compute bound, and makes three recommendations to fully utilize the GPUs' potential for relational workloads when the memory wall is removed.

S. Hepkema, Bo-Wen Wu, Christos Kozyrakis et al. · 0 citations
Preprint Aug 2026

Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study

Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. It is common for a single compute node (intranode) to host multiple GPUs, typically four or more. Therefore, effectively leveraging all these GPUs within a s...

E. Krishnasamy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.