This work presents GreenPipe, an automated profiling-training-validation pipeline that builds multi-resource regression models from external power meter measurements and attributes power to containers proportionally, exposing performance-energy trade-offs across workload configurations.
Abstract
Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network. Accurate online power estimation on resource-constrained ARM nodes without hardware power counters such as RAPL remains a challenge, and CPU-only models fail to capture multi-resource behavior. We present GreenPipe, an automated profiling-training-validation pipeline that builds multi-resource regression models from external power meter measurements and attributes power to containers proportionally. GreenPipe is evaluated on a Raspberry Pi 4 edge node in a K3s edge-cloud testbed, covering DNN inference with three vision models, multiple precisions, thread counts, and both local and serving scenarios. System-level MAPE is 6.3-9.4%, improving over CPU-stress and utilization-only baselines by 26.9% MAPE on average. We jointly report inference latency and energy per inference, exposing performance-energy trade-offs across workload configurations.
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactiv...
Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al.· 0 citations
Hydra is presented, a common-schema, phase-aware workload characterization framework for LLM inference on edge SoCs that enables reproducible, phase-aware characterization of edge LLM inference.
Amir Taherin, Sana Taghipour Anvari, Charles Amante et al.· 2 citations
PRISM, a prediction-guided runtime framework that jointly selects model variants and CPU allocations for containerized edge microservices, and adapts each pipeline stage in place and minimizes predicted CPU-package energy under deadline, resource, and offline model-level Quality of Result constraints is presented.
Uwe Gropengießer, Thomas Reuter, Dominik Schön et al.· 0 citations
An integrated framework that combines LLM fine-tuning using federated adaptive local low-rank adaptation (FedALoRA) with forecasting-driven autoscaling policies is proposed, making it well-suited for federated-LLM Autoscaling in dynamic edge environments.
With its elastic power and a pay-as-you-go cost model, the deployment of deep learning inference services (DLISs) on serverless platforms is emerging as a prevalent trend. However, the varying resource requirements of different layers in DL models hinder resource utilization and increase costs, when DLISs are deployed...
Ji-Aang Duan, Shiyou Qian, Dingyu Yang et al.· Proceedings of the Internati...· 7 citations
A carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware and matches cloud-level accuracy while reducing operational carbon emissions by $4\times on average.
Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, S. Tragoudas et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.