Skip to content

On the Limits of Machine-Learned Ranking for Modern Microarchitectural Policies

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

It is shown that high cycle or aggregate ranking accuracy can reflect mastery of easy, high-margin cases while missing the local reversals that carry the most architectural insight and for which cycle-level simulation remains indispensable.

Abstract

Machine-learning predictors estimate processor performance far faster than cycle-level simulation. For design-space exploration, however, the valuable test is not merely reproducing the usual hardware ordering, but identifying how different hardware configurations rank on individual program phases. We evaluate four ML-predictors in two design regimes: \emph{Structural Parameters} (SP), varying hardware resources such as issue width, ROB size, and cache capacity; and \emph{Behavioral Policies} (BP), varying prefetching and replacement algorithms. In the SP regime, aggregate ranking is strong, yet counter-intuitive windows(CIW)---where the configuration expected to be slower is faster---constitute $22.4\%$ of non-tied windows across five pairs with a clear architectural prior. CIW match across these pairs is only $23.3$--$39.9\%$; every point estimate is below the $50\%$ random strict-ordering reference. The BP regime presents a different failure: ground-truth ties cover $37.8\%$ of pair-windows, most strict pairs have margins of only a few cycles, and no model family reliably beats a feature-free majority baseline. NeuroScalar and SimNet fall below that baseline, Concorde is statistically tied with it, and the best selected OneDSE head improves by only $2.1$ percentage points. Accuracy rises mainly at large margins. We further show that this failure is not a matter of model capacity: an information-theoretic analysis reveals that when ranking outcomes depend on hidden microarchitectural state absent from the instruction stream, no trace-based predictor can exceed the Bayes accuracy determined by observable inputs alone. Thus high cycle or aggregate ranking accuracy can reflect mastery of easy, high-margin cases while missing the local reversals that carry the most architectural insight and for which cycle-level simulation remains indispensable.

View source

Similar papers

Preprint Aug 2026

Aneto: Predicting System Performance by Exploiting Cross-Workload Regularity

Aneto is a mechanistic-empirical regression model that estimates the performance-latency sensitivity of any new workload from a single run, enabling first-order CPI prediction under any memory configuration.

Raúl Taranco, Rene Mueller, Michael Giardino · 0 citations
Book Open access Sep 2026

Code Generation from Regression Trees for Microsecond-Scale Decisions in Operating Systems

In today's world of heterogeneous server hardware, deciding on suitable task and data placements is a far from trivial undertaking. Depending on system load and application behaviour, some compute and memory assignments improve performance, while others impair it. Yet, these increasingly complex decisions must be made...

B. Friesel, Marcel Lütke Dreimann, Olaf Spinczyk · 0 citations
Preprint Sep 2026

Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the Predictor

Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy. We disentangle these variables with matched controls: the same admission gate is applied to both a 257-parameter online MLP and a classical stride predic...

Youssef Majdane, Simone Jarno Casartelli, Enrico Lopedoto · 0 citations
#artificial intelligence Preprint Sep 2026

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical...

Da Zhao, K. Sankaralingam, Christos Kozyrakis et al. · 0 citations
Preprint Aug 2026

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?

Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanoseconds of latency determines competitive advantage and designs iterate continuously as protocols, strategies, and regulations evolve. FinHardBench, a benchmark of 33 fin...

Weimin Fu, He-Jia Zhang, Ming-Hao Shao et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.