PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases and achieves a lower energy-delay product than the existing governor across all evaluated workloads.
Abstract
Existing mobile GPU dynamic voltage and frequency scaling (DVFS) policies rely on coarse-grained utilization metrics and treat the GPU as an isolated control domain, failing to reflect the layer-level computational and memory diversity of deep neural network (DNN) inference. This paper proposes PUMA, a performance monitoring unit (PMU)-guided multi-domain layer-aware DVFS framework. PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases. In the offline stage, representative DNN workloads were profiled across 504 frequency combinations spanning the GPU, memory interface (MIF), and internal interconnect (INT) domains to derive PMU thresholds and domain-specific frequency-correction rules. At runtime, PUMA applies threshold-based bounded corrections to the GPU, MIF, and INT frequency decisions on top of the existing governors. PUMA was implemented at the kernel level on Google Pixel 9 and evaluated using six DNN workloads. Compared with the existing governor, PUMA reduced SoC power by 28.54% on average and by up to 33.41%, while reducing energy by 23.76% on average and by up to 27.97%, with an average inference latency increase of 7.28% and a maximum increase of 12.82%. Compared with GPU-only correction, full PUMA further reduces average power by 11.4% relative to GPU-only correction, and PUMA achieves a lower energy-delay product than the existing governor across all evaluated workloads.
SA-Scheduler is presented, a structure-derived bottleneck-aware scheduling framework for multitasking MCM-GPUs that determines chip placement without hardware modification or runtime bottleneck profiling, and provides a principled and scalable foundation for multitasking on future MCM-GPUs.
Tie-Jian Zhang, Guangda Zhang, Lu Wang et al.· ACM Transactions on Design A...· 0 citations
This work evaluates whether Model FLOPs Utilization (MFU) can serve as a portable, software-defined predictor of GPU power for LLMs, finding that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training.
This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the underlying acceleration mechanism. The design accelerates PDSCH, PUSCH, SRS, PRACH, split-8 lower-PHY transforms, and O-RAN...
M. Pennybacker, Wanze Liu, A. Kharchenko et al.· 2 citations
TierKV is presented, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO), which improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, wh...
Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al.· 0 citations
The authors co-design WIFiV-LPDDR, a wide-I/O LPDDR system for precision-tagged reads, BRQ-KV for a canonical low-rank-plus-INT8-residual prefix with query-dependent per-entry precision, and DAT-FFN for drift-mapped canonical replacement, adjacent-stage-corrected low-bit delta, or cached-state carry while keeping live...
Wei-Hsing Huang, Kiseok Lee, Ming-Yen Lee et al.· 0 citations
Evaluation using cycle-accurate hardware-software co-simulation demonstrates that CHIPSMORE effectively sustains high throughput across varying model sizes, context lengths, and batch sizes while maintaining favorable power scaling.