Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.
Cun-Chen Hu, Liangliang Xu, Tianyu Liu et al.· 0 citations
UVirtio introduces a device-profile-based virtual hardware abstraction layer that minimizes performance overhead, and implements a live migration mechanism using differential packing, providing a scalable and agile virtualization solution for the ubiquitous computing frontier.
Muliang Shou, Yufan Jiang, Tianlei Xiong et al.· ACM Transactions on Architec...· 0 citations
LEVELLER is proposed, the first communication scheduling system that achieves max-min fairness specifically for DLT workloads and introduces a novel online metric, normalized progress rate, which quantifies training experience by measuring actual progress against a contention-free ideal.
Geng Li, Yang Li, Mingyuan Zang et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.