Skip to content

Energy-Efficient GPU DVFS for Fine-Tuning of SLMs on Resource-constrained Embedded Devices

Jul 2026 · arXiv.org · Vol abs/2607.05933 · 0 citations · 19 references
Computer Science

TL;DR

This work characterize the fine-tuning behavior of representative encoder-only SLMs of BERT variants, and autoregressive decoder-only SLMs of Pythia variants on GLUE benchmarks, and proposes a simple yet effective ML-based model selection that selects energy-optimal GPU DVFS settings on resource-constrained embedded platforms.

Abstract

Dynamic Voltage Frequency Scaling (DVFS) on resource-constrained embedded GPU platforms is essential for energy-efficient small language model (SLM) fine-tuning, as privacy- and personalization-driven adaptation increasingly requires local execution and involves repeated forward-backward optimization over many mini-batches, making it substantially more time- and energy-intensive than single-pass inference. To this end, 1) we first characterize the fine-tuning behavior of representative encoder-only SLMs of BERT variants, and autoregressive decoder-only SLMs of Pythia variants on GLUE benchmarks. In addition to the characterizations, 2) we propose a simple yet effective ML-based model selection that selects energy-optimal GPU DVFS settings on resource-constrained embedded platforms. Our results on NVIDIA Jetson AGX Orin demonstrate average 13.11% energy savings (up to 26.73%) over MAXN Mode 0, which has no explicit power cap.

View source

Similar papers

Preprint Aug 2026

Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training

This work evaluates whether Model FLOPs Utilization (MFU) can serve as a portable, software-defined predictor of GPU power for LLMs, finding that a linear MFU-based power model fits every tested GPU as long as the workload is compute-bound, as in production LLM training.

Niklas Enskat, Philipp Wiesner · 1 citation · ⚡1
Open access 2026

PUMA: A PMU-Guided Multi-Domain Layer-Aware DVFS Framework for Low-Power Mobile AI on Smartphones

PUMA combines offline layer/layer-chain PMU characterization with online GPU PMU observations to identify execution characteristics associated with compute-bound, memory-bound, and bursty phases and achieves a lower energy-delay product than the existing governor across all evaluated workloads.

W. Chang, Seung-Ryeol Ohk, Young-Jin Kim · 0 citations
#artificial intelligence Preprint Aug 2026

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

Faster Flash Decoding (FFD) is presented, a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding and introduces the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization.

Zhigeng Liu, Zhiyuan Ning, Rui-Xiao Li et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.