Skip to content
Preprint

PANEM: A Heuristic Latency Model

Aug 2026 · 1 citation · 26 references
Computer Science

TL;DR

PANEM is presented, a lightweight event-driven heuristic model that has guided four generations of commercial server-core development at Ampere Computing and shows that a calibrated, contention-aware abstraction can deliver practical predictive value for industrial design-space exploration at simulation costs similar to fixed-latency models.

Abstract

Accurate pre-silicon memory modeling is essential for achieving meaningful representation of workloads on cloud-class many-core processors. Existing options force a poor tradeoff between fidelity and speed: fixed-latency models are fast but misleading, while cycle-accurate DRAM models are costly and difficult to scale across large study spaces or onto single-core environments. This paper presents PANEM, a lightweight event-driven heuristic model that has guided four generations of commercial server-core development at Ampere Computing. PANEM converts bandwidth-latency characterization data into a dynamic request-bytes/latency response, allowing miss latency to adapt to transient demand, queuing pressure, and read/write mix during simulation. Integrated into a single-core flow with configurable system-loading assumptions, PANEM enables realistic bandwidth constraints and contention-aware latency behavior without sacrificing throughput. Across a broad cloud workload trace suite, PANEM avoids the optimistic and pessimistic biases of fixed-latency baselines, yields more reliable conclusions for prefetching and dynamic throttling studies, and materially improves core-resource sizing decisions. These results show that a calibrated, contention-aware abstraction can deliver practical predictive value for industrial design-space exploration at simulation costs similar to fixed-latency models.

View source

Similar papers

Preprint Aug 2026

Aneto: Predicting System Performance by Exploiting Cross-Workload Regularity

Aneto is a mechanistic-empirical regression model that estimates the performance-latency sensitivity of any new workload from a single run, enabling first-order CPI prediction under any memory configuration.

Raúl Taranco, Rene Mueller, Michael Giardino · 0 citations
Open access Aug 2026

A Cache Modeling Framework for Accelerator Systems

Modern hardware accelerators increasingly rely on cache-based memory systems to improve modularity and tolerate irregular memory behavior in sparse and data-dependent workloads. However, configuring accelerator caches remains challenging: designers must choose cache capacity, Miss Status Holding Register (MSHR) count,...

Lingfeng Pei, Wei Siew Liew, Udaree Kanewala et al. · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al. · 0 citations
Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.