From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs
This report presents an analytically structured, empirically calibrated, GPU-level methodology for estimating LLM inference energy on NVIDIA H100-class accelerators without direct runtime measurement, and provides transparent, reproducible, and assumption-explicit approximations suitable for model comparison, green-cod...