Aug 2026· 1 citation· ⚡ 1 influential· 31 references
Computer Science
TL;DR
This work evaluates whether Model FLOPs Utilization (MFU) can serve as a portable, software-defined predictor of GPU power for LLMs, finding that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training.
Abstract
High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption. Established GPU power models rely on hardware utilization counters, which do not exist until the workload has actually run. This work evaluates whether Model FLOPs Utilization (MFU)-an analytical, software-defined metric relating achieved throughput to peak hardware capability-can serve as a portable, software-defined predictor of GPU power for LLMs. We benchmark almost 3000 single-device training runs across six GPUs, covering different model families, numerical precisions, batch sizes, and context-window lengths. We find that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training. Fitting per-(GPU, dtype, batch size) instead of per-GPU drops the within-cell mean error from around 10% to around 1%, matching the cross-repeat measurement-noise floor.
The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground...
Lukas Groth, Andrija Nešković, Rainer Buchty et al.· ACM Transactions on Embedded...· 0 citations
Graphics Processing Units (GPUs) have been serving as critical computation resources for large-scale parallel computations. With increasing chip complexity, power efficiency has become an important design objective for modern GPUs. GPU power optimization relies on fast power evaluation, requiring architecture-level GPU...
Qi-Jun Zhang, Yao Lu, Shang Liu et al.· 0 citations
The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....
Matthew Tindale, I. Karlin, Tobias Salamon et al.· Inquiry@Queen's Undergraduat...· 0 citations
SLIM (Saturation-Aware Lightweight Performance Model), a semi-analytical model that predicts LLM inference throughput and latency from analytical formulations of Transformer computation and memory traffic, is introduced, which outperforms representative performance-modeling baselines while successfully generalizing to...
Pol G. Recasens, F. Agulló, Yue Zhu et al.· arXiv.org· 0 citations
Three fundamental design principles are revealed that provide design-space guidance for architects designing the next generation of memory-accelerated LLM systems.
Corey Lammie, Hadjer Benmeziane, W. Simon et al.· 0 citations
A mathematical model is proposed to predict throughput and energy consumption for concurrently executing CV workloads on edge GPU accelerators and can be integrated into functional simulation frameworks for edge–cloud deployment studies.
Abhinaba Chakraborty, D. Colle, M. Pickavet et al.· Journal of Real-Time Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.