Skip to content
Preprint

Themis: Software-Defined Hardware Prefetching

Jul 2026 · 0 citations · 116 references
Computer Science

TL;DR

Themis is a profile-guided hardware prefetching solution that implements a novel hardware-software interface for data prefetching: the software directs the hardware on where to prefetch, and the hardware identifies and issues prefetches in the regions of interest.

Abstract

Data cache misses represent a significant portion of stall cycles in datacenter workloads. Hardware prefetchers that reduce such stalls by fetching data ahead of time have become increasingly sophisticated. However, to achieve high coverage, they have to prefetch aggressively, generating many inaccurate accesses that waste memory bandwidth. This is problematic in datacenter environments where memory bandwidth is a limited resource due to high multi-tenancy. We observe that for datacenter workloads, inaccurate prefetches can be effectively filtered on a data page granularity, without sacrificing prefetch coverage. However, storing per-page metadata about prefetch usefulness in hardware is costly, so we propose a novel hardware-software interface for data prefetching: The software directs the hardware on where to prefetch, and the hardware identifies and issues prefetches in the regions of interest. We propose Themis, a profile-guided hardware prefetching solution that implements this new interface. Themis utilizes page-level hints stored in page-table entries to disable the prefetcher for certain data pages at runtime. Themis requires no binary or ISA changes and can be used to optimize processes without disrupting their execution. Themis is also orthogonal to existing works on prefetching and can be applied to optimize any hardware prefetcher. Our results show that Themis is able to achieve around 40% reduction in useless prefetch requests, resulting in speedup for all the evaluated prefetchers for datacenter workloads, including 4.1% for BOP, 3.1% for SPP+PPF, and 1.4% for Pythia.

View source

Similar papers

Book Open access Sep 2026

BASIC-Prefetcher: Bin-based Address and Size-Informed Caching for AI-Driven SSD Workloads

Read latency is a critical bottleneck for NAND-based SSDs in AI-driven datacenter workloads, where model parameters and key-value data are frequently swapped between main memory and storage. Existing prefetching schemes operate on block-level address sequences that have been stripped of application context by the files...

Han Jang, Dongjun Lee, Youngbin Jin et al. · 0 citations
Preprint Aug 2026

Why Do Prefetchers Fail? Let Agents Answer

To the authors' knowledge, this is the first empirical demonstration that an agent-driven hardware-design process can produce an RTL-practical prefetcher that outperforms state-of-the-art human designs on unseen workloads.

Xiang-Feng Sun, Ce-Yu Xu, Ningzhi Ai et al. · 0 citations
Conference Aug 2026

Paging-Resilient Prefetching in Flash-Based CXL SSDs

CXL SSDs extend system memory using NAND flash, providing a scalable solution to the capacity and bandwidth limits of memory-intensive, multi-tenant cloud services. Since CXL SSDs are directly addressable by the CPU, SSD-internal prefetching is crucial for hiding flash-grade latency from the host. However, host-side pa...

Chung-Min Yu, Chih-Kang Yeh, Ying-Shuo Lin et al. · 0 citations
Open access Aug 2026

Machine Learning-Driven Optimization of Cache Memory Prefetching Processes

The proposed model enhances cache prefetching by implementing an LSTM-based prefetcher that learns from dynamic program traces, thereby eliminating the linear relationship between fetch count and space while enhancing the capability to identify and forecast intricate access patterns.

Remegius Praveen Sahayaraj L, A. E, Aswini E · 0 citations
Open access Sep 2026

Packets are Not Pages: Flow-Based Addressing Conserves Memory Bandwidth

SmartNICs promise hardware offloading for network applications like traffic analysis and virtual host dispatching. However, due to memory bandwidth limitations, SmartNICs are unable to effectively accelerate applications like intrusion detection and deep packet inspection that require high-speed reassembly. We argue th...

Agur Adams, H. Shim, Colin Drewes et al. · 0 citations
Open access Aug 2026

SAI: Virtualizing Shared Memory of GPU for AI workload acceleration

This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...

Hanqing Li, Tie-Jun Li, Sheng Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.