Skip to content

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

Jul 2026 · arXiv.org · Vol abs/2607.26571 · 1 citation · 25 references
Computer Science

TL;DR

This report presents an analytically structured, empirically calibrated, GPU-level methodology for estimating LLM inference energy on NVIDIA H100-class accelerators without direct runtime measurement, and provides transparent, reproducible, and assumption-explicit approximations suitable for model comparison, green-coding analysis, and design-time evaluation of LLM inference workloads.

Abstract

The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprint of deployed AI systems. However, direct measurement of inference energy often requires hardware telemetry, power instrumentation, or infrastructure-specific monitoring, limiting its applicability in comparative studies, early-stage system design, and sustainability reporting. This report presents an analytically structured, empirically calibrated, GPU-level methodology for estimating LLM inference energy on NVIDIA H100-class accelerators without direct runtime measurement. The proposed estimator combines parameter-scaled transformer FLOP accounting, calibrated memory-traffic factors, and hardware-specific energy coefficients for FP16/BF16 tensor-core computation and high-bandwidth-memory movement. It explicitly separates prompt prefill from autoregressive decoding, enabling energy estimates for input tokens, output tokens, and complete inference requests. The methodology further decomposes total energy into compute, parameter-access, key-value-cache write, and attention-read components, allowing the scaling behavior with model size, context length, and generated-token count to be analyzed. The resulting estimates are not intended to replace physical power measurements; rather, they provide transparent, reproducible, and assumption-explicit approximations suitable for model comparison, green-coding analysis, and design-time evaluation of LLM inference workloads.

View source

Similar papers

Book Open access Jul 2026

Phase-Wise Analysis of LLM Inference Acceleration on GPU, CPU, and Edge Device

This study presents a cross-platform, multi-model empirical study, where several important observations are brought, including the contrastive effect of quantization under different hardware bottlenecks, along with a quantification of runtime delays caused by the lack of parallelism in the ARM architecture.

Subhransu Das, Jiaming Cheng, Swathi Vallabhajosyula et al. · 0 citations
Conference Jul 2026

A Machine Learning Approach to Estimating Energy Use in Language Model Inference

The rapid expansion of Large Language Model (LLM) serving in cloud data centers has created a critical need for energy-aware scheduling. However, estimating inference energy typically requires hardware-level power telemetry, which is rarely accessible to cloud tenants. This paper proposes a lightweight, machine-learnin...

Bediga Sharan, Swarup Ghosh · 0 citations
#artificial intelligence Preprint Sep 2026

A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactiv...

Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al. · 0 citations
Preprint Aug 2026

Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode

This report argues that the most effective response to single-token autoregressive decode on CPUs is to co-design the model architecture and the inference runtime together, and presents cflow, a CPU-first streaming engine, alongside a family of pipeline-native transformer architectures whose inter-layer dependency grap...

Tom Poperszky · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.