Into the CUDA Multiverse of Runtime Compilation: Exploring Kernel Fusion Strategies for GPU Database Systems
Abstract
This demonstration presents the CUDA Code-General, a query compilation framework for GPU database systems. The interface allows selecting any of the 22 TPC-H benchmark queries or entering arbitrary SQL, to choose between runtime or static compilation and to inspect the generated CUDA code for three kernel fusion strategies side by side: fused multiple sequential kernels, cooperative groups, and dynamic parallelism. Users can observe how the number of generated pipelines affects compilation cost, how cooperative groups fuse them into a single kernel with synchronisation barriers, and how dynamic parallelism delegates kernel launches to the device.