Skip to content
Open access

Compression-Aware GPU Buffer Management

Sep 2026 · Datenbank-Spektrum · 0 citations · 29 references

TL;DR

A three-tier buffer manager that manages encoded pages across SSD, RAM and GPU device memory while codec-specific operators fuse decompression, selection, join and aggregation is presented.

Abstract

Hardware accelerators, such as graphics processing units (GPUs) connected via PCIe, provide their own byte-addressable device memory. When integrating their memory into the global buffer pool, database systems should place compressed data and the associated operators on the best suited units to maximise overall throughput and bandwidth. This paper presents a three-tier buffer manager that manages encoded pages across SSD, RAM and GPU device memory while codec-specific operators fuse decompression, selection, join and aggregation. Splitting the workload between CPU (uncompressed in RAM) and GPU (compressed on device) and combining partial aggregates yields a throughput of 3.76 GiB/s, which is 15% higher than GPU-only execution. The results are the basis for a cost-model for data and operator placement.

Read PDF

Similar papers

Book Open access Sep 2026

Analysis of shared memory between CPUs and GPUs

This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.

Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al. · 0 citations
Conference Sep 2026

MPTAC: Memory Controller Partitioning and Traffic-Aware Contention Control in GPUs

The advent of cloud-based artificial intelligence and the increased digitalization of embedded systems require powerful GPUs capable of simultaneously running kernels from different software providers. To accommodate the resource isolation and Execution Time Determinism (ETD) needed with the increasing number of kernel...

Vahid Geraeinejad, Paul Delestrac, Javier Barrera et al. · 0 citations
Open access Sep 2026

G-CasDec: General Cascaded Decompression on GPUs

GPU-accelerated analytical query processing is often limited by both GPU device memory capacity and host-to-device data transfer time. Modern data compression techniques, such as cascaded lightweight compression, can mitigate these issues. However, existing designs all exhibit critical tradeoffs on compression ratios,...

Yong-Qi Zhuo, Xin-Yu Zeng, Huan-Chen Zhang et al. · 0 citations
Book Open access Aug 2026

A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU

This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.

Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah et al. · 0 citations

FitFloat: Read/Write Random-Access Compressed Floating-Point Arrays for GPUs

FitFloat is presented, a drop-in floating-point array replacement supporting user-specified precision on GPUs with the goal of reducing storage requirements of scientific applications while maximizing performance over Unified Memory.

Andrew Rodriguez, Martin Burtscher · 0 citations
Book Open access Sep 2026

Bridging CPU and GPU I/O with a Device-Resident File System

Modern AI and data analytics workloads have terabyte-scale datasets, making GPUs access storage frequently. GPUDirect Storage allows the data path from SSD to GPU to bypass CPU memory. However, the file system shared by the CPU and GPU still runs on the CPU. Existing GPU file systems cache metadata to avoid frequent ac...

Qi Chen, Guan-Yi Chen, Yu-Jie Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.