It is shown that semantic understanding and reasoning about a program is a vital component in inserting effective software prefetches, and that prefetching at the scale of large codebases poses new challenges, including prefetch codependence and interference.
Abstract
Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic knowledge of the workload at hand. However, inserting software prefetches is a time consuming, unpredictable task, with a high degree of sensitivity to microarchitecture or future code changes. While several compiler- and FDO-based approaches insert and tune software prefetches automatically, they use heuristics that do not generalize and do not consider issues that arise in large codebases. In this paper, we show that semantic understanding and reasoning about a program is a vital component in inserting effective software prefetches. We also identify that prefetching at the scale of large codebases poses new challenges, including prefetch codependence and interference. To remedy this, we build Presage, a system that leverages the semantic reasoning of Large Language Models to insert effective software prefetches. Presage processes a workload through a custom agent harness specifically designed to navigate the wide space of potential prefetches on large-scale workloads. First, a proposer agent with access to performance metrics identifies promising software prefetching candidates. Then, several long-horizon optimizer agents optimize prefetching candidates in a loop, working towards discovery of performance wins. Finally, a combiner agent handles composition of individually effective prefetches. Through this method, Presage is able to insert prefetches that improve performance by a geomean of 10% across 81 workloads by exploring tens of different prefetching alternatives per workload, including a 2.7% runtime reduction on SPEC CPU 2026 where prior SOTA fails. When compared against prior SOTA on its evaluation suite, Presage achieves a geomean runtime improvement of 17% compared to SOTA's 10%.
This work presents ArchAgent v2, a framework which scales automated microarchitecture search to multi-level data prefetching and introduces two new additions to ArchAgent: a cascaded evolutionary search that subdivides the design space by sequentially evolving and freezing prefetchers at individual cache levels, and a...
Abraham Gonzalez, Raghav Gupta, Akanksha Jain et al.· 0 citations
To the authors' knowledge, this is the first empirical demonstration that an agent-driven hardware-design process can produce an RTL-practical prefetcher that outperforms state-of-the-art human designs on unseen workloads.
Xiang-Feng Sun, Ce-Yu Xu, Ningzhi Ai et al.· 0 citations
The proposed model enhances cache prefetching by implementing an LSTM-based prefetcher that learns from dynamic program traces, thereby eliminating the linear relationship between fetch count and space while enhancing the capability to identify and forecast intricate access patterns.
Remegius Praveen Sahayaraj L, A. E, Aswini E· international journal of eng...· 0 citations
Modern web applications demand sustained low latency under workloads that shift across users, devices, sessions, and network conditions. Classical cache replacement policies such as Least Recently Used (LRU) and Least Frequently Used (LFU) treat every cached object identically and ignore the cost-of-miss heterogeneity...
Akshatha Madapura Anantharamu· World Journal of Advanced En...· 0 citations
Database management systems increasingly serve dynamic and exploratory workloads, yet many of their decisions still rely on low-level signals such as recency, frequency, and address locality. These signals capture how data was accessed, but not what is being examined or how an analytical focus evolves. We argue for tre...
Farzaneh Zirak, Kasper Overgaard Mortensen, F. Choudhury et al.· 0 citations
Caching is widely used across the system stack to improve performance and efficiency, with eviction algorithms at its core. Existing cache eviction policies fall into two broad categories: static heuristics (e.g., 2Q, S3-FIFO) and smart algorithms (e.g., ARC, LRB). Smart caches can adapt to workloads and have the poten...
Haocheng Xia, William Nixon, Bintang Dwi Marthen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.