Skip to content
#edge computing Book Open access

Pagoda: Time and Energy Rooflines for DNN Workloads on Edge Accelerators Across Power Modes

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 11 references

TL;DR

This work develops time and energy roofline characterizations for Orin AGX across power modes, coupled with analytical models of compute and memory for DNN inference, and extends these models to DNN training, and demonstrates power-mode tuning that achieves up to \(15\%\) lower energy with small inference-time impact for some workloads.

Abstract

Edge accelerators such as Nvidia Jetsons are an integral part of the computing continuum, and are used for DNN inference and training. Jetson edge devices have 2000+ CUDA cores within a 70W power envelope and offer 1000s of power modes to customize CPU, GPU and memory frequencies, enabling diverse power–performance trade-offs for energy-constrained deployments. While data-driven methods exist to predict power and latency, there is limited principled understanding of why different power modes perform the way they do. We develop time and energy roofline characterizations for Orin AGX across power modes, coupled with analytical models of compute and memory for DNN inference. These reveal quantitative insights, e.g., the default power mode is not the most energy-efficient and time efficiency implies energy efficiency. We extend our models to DNN training, and demonstrate power-mode tuning that achieves up to \(15\%\) lower energy with small inference-time impact for some workloads.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment n...

Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al. · 0 citations
#artificial intelligence Conference Open access Aug 2026

Greenbench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source Llm Inference on Apple Silicon

GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.

R. Kannan, Rajendra P. Firke, Shreya Bengle et al. · 0 citations
Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-i...

Siddharth Patel, Rohit Singh · 0 citations
Aug 2026

Modeling throughput and power consumption for real-time concurrent vision application deployments on edge

A mathematical model is proposed to predict throughput and energy consumption for concurrently executing CV workloads on edge GPU accelerators and can be integrated into functional simulation frameworks for edge–cloud deployment studies.

Abhinaba Chakraborty, D. Colle, M. Pickavet et al. · 0 citations
Open access Aug 2026

Uncovering Power and Frequency Behaviors and Their Security Implications in Edge GPU Platforms

Edge platforms increasingly rely on GPUs and deep learning accelerators (DLAs) to support real-time computing of emerging workloads. However, their power, frequency, and security characteristics remain largely unexplored. This paper provides the first in-depth characterization of power and frequency behaviors of NVIDIA...

M. Rafi, Kevin Chau, Hyeran Jeon · 0 citations
#edge computing Sep 2026

Compatibility Ratio as a Guideline for Hardware/DNN Co-Design of Embedded Accelerators

The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground...

Lukas Groth, Andrija Nešković, Rainer Buchty et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.