RTX 5080 LLM Power Efficiency: Measured Watts and Joules per Token
Measured board power, tokens per joule, and electricity cost per million tokens for local large language model inference on a retail NVIDIA RTX 5080. Board power was logged with nvidia-smi at 1 Hz while driving fixed-length generations. Efficiency. Llama 3.2 3B 1.20 tokens/joule; sparse gpt-oss 20B 0.665; Qwen 2.5 7B 0.53; Qwen 2.5 14B 0.317. The sparse 20.9B model is approximately twice as efficient per joule as the dense 14B, indicating that architecture and quantization dominate parameter count on the efficiency axis. Loaded power. Board power medians of 265-344 W against a 448 W stock limit that was never reached, at 82-93% utilization and 49-52 degrees C. Idle finding. Across a week of captures the card idled at 52-71 W at the Windows desktop, pinned in P0 with graphics clocks near 2.9 GHz; the cleanest achievable state still read 53.7 W. Published review figures typically quote single-digit to 15 W idle. At the May 2026 EIA US residential average of $0.184/kWh, 52-71 W continuous is 456-622 kWh, or approximately $85-115 per year before any tokens are generated. Since generation itself costs only $0.04-$0.16 per million tokens, idle behaviour rather than model choice dominates the operating cost of an intermittently used inference node. Includes raw 1 Hz telemetry captures in addition to summary rows. Canonical page, full method and change log: https://techfuelhq.com/data/rtx-5080-llm-power-efficiency/