Sep 2026· Proceedings of the International Conference on Parallel Processing· 0 citations· 11 references
TL;DR
This work develops time and energy roofline characterizations for Orin AGX across power modes, coupled with analytical models of compute and memory for DNN inference, and extends these models to DNN training, and demonstrates power-mode tuning that achieves up to \(15\%\) lower energy with small inference-time impact for some workloads.
Abstract
Edge accelerators such as Nvidia Jetsons are an integral part of the computing continuum, and are used for DNN inference and training. Jetson edge devices have 2000+ CUDA cores within a 70W power envelope and offer 1000s of power modes to customize CPU, GPU and memory frequencies, enabling diverse power–performance trade-offs for energy-constrained deployments. While data-driven methods exist to predict power and latency, there is limited principled understanding of why different power modes perform the way they do. We develop time and energy roofline characterizations for Orin AGX across power modes, coupled with analytical models of compute and memory for DNN inference. These reveal quantitative insights, e.g., the default power mode is not the most energy-efficient and time efficiency implies energy efficiency. We extend our models to DNN training, and demonstrate power-mode tuning that achieves up to \(15\%\) lower energy with small inference-time impact for some workloads.
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment n...
Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al.· 0 citations
GreenBench, a benchmarking framework that evaluates the energy efficiency, throughput, and carbon footprint of five open-source LLMs across three NLP tasks on an Apple M4 Pro with 48 GB unified memory, is presented.
R. Kannan, Rajendra P. Firke, Shreya Bengle et al.· International Conference on...· 0 citations
Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-i...
A mathematical model is proposed to predict throughput and energy consumption for concurrently executing CV workloads on edge GPU accelerators and can be integrated into functional simulation frameworks for edge–cloud deployment studies.
Abhinaba Chakraborty, D. Colle, M. Pickavet et al.· Journal of Real-Time Image P...· 0 citations
Edge platforms increasingly rely on GPUs and deep learning accelerators (DLAs) to support real-time computing of emerging workloads. However, their power, frequency, and security characteristics remain largely unexplored. This paper provides the first in-depth characterization of power and frequency behaviors of NVIDIA...
M. Rafi, Kevin Chau, Hyeran Jeon· ACM Transactions on Architec...· 0 citations
The Compatibility Ratio (CR) is introduced as a simple guideline for evaluating performance trade-offs between optimal hardware micro-architecture configurations across different workloads and shows that, for the considered accelerator, a DNN model-family optimized configuration might occupy an effective middle ground...
Lukas Groth, Andrija Nešković, Rainer Buchty et al.· ACM Transactions on Embedded...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…