Skip to content
Open access

PUMA: A PMU-Guided Multi-Domain Layer-Aware DVFS Framework for Low-Power Mobile AI on Smartphones

2026 · IEEE Access · Vol 14, pp. 130512-130525 · 0 citations · 42 references

TL;DR

PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases and achieves a lower energy-delay product than the existing governor across all evaluated workloads.

Abstract

Existing mobile GPU dynamic voltage and frequency scaling (DVFS) policies rely on coarse-grained utilization metrics and treat the GPU as an isolated control domain, failing to reflect the layer-level computational and memory diversity of deep neural network (DNN) inference. This paper proposes PUMA, a performance monitoring unit (PMU)-guided multi-domain layer-aware DVFS framework. PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases. In the offline stage, representative DNN workloads were profiled across 504 frequency combinations spanning the GPU, memory interface (MIF), and internal interconnect (INT) domains to derive PMU thresholds and domain-specific frequency-correction rules. At runtime, PUMA applies threshold-based bounded corrections to the GPU, MIF, and INT frequency decisions on top of the existing governors. PUMA was implemented at the kernel level on Google Pixel 9 and evaluated using six DNN workloads. Compared with the existing governor, PUMA reduced SoC power by 28.54% on average and by up to 33.41%, while reducing energy by 23.76% on average and by up to 27.97%, with an average inference latency increase of 7.28% and a maximum increase of 12.82%. Compared with GPU-only correction, full PUMA further reduces average power by 11.4% relative to GPU-only correction, and PUMA achieves a lower energy-delay product than the existing governor across all evaluated workloads.

Read PDF

Similar papers

Open access Aug 2026

Structure-Derived Bottleneck-Aware Scheduling for Multitasking MCM-GPUs

SA-Scheduler is presented, a structure-derived bottleneck-aware scheduling framework for multitasking MCM-GPUs that determines chip placement without hardware modification or runtime bottleneck profiling, and provides a principled and scalable foundation for multitasking on future MCM-GPUs.

Tie-Jian Zhang, Guangda Zhang, Lu Wang et al. · 0 citations
Preprint Aug 2026

Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training

This work evaluates whether Model FLOPs Utilization (MFU) can serve as a portable, software-defined predictor of GPU power for LLMs, finding that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training.

Niklas Enskat, Philipp Wiesner · 1 citation · ⚡1
Preprint Aug 2026

GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance

This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the underlying acceleration mechanism. The design accelerates PDSCH, PUSCH, SRS, PRACH, split-8 lower-PHY transforms, and O-RAN...

M. Pennybacker, Wanze Liu, A. Kharchenko et al. · 2 citations
#machine learning Preprint Sep 2026

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...

Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al. · 0 citations
Preprint Sep 2026

Hardware Acceleration of Block-Diffusion LLM for Edge Devices

The authors co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live...

Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.