Skip to content
Preprint

GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge Nodes

Sep 2026 · 0 citations · 12 references
Computer Science

TL;DR

This work presents GreenPipe, an automated profiling-training-validation pipeline that builds multi-resource regression models from external power meter measurements and attributes power to containers proportionally, exposing performance-energy trade-offs across workload configurations.

Abstract

Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network. Accurate online power estimation on resource-constrained ARM nodes without hardware power counters such as RAPL remains a challenge, and CPU-only models fail to capture multi-resource behavior. We present GreenPipe, an automated profiling-training-validation pipeline that builds multi-resource regression models from external power meter measurements and attributes power to containers proportionally. GreenPipe is evaluated on a Raspberry Pi 4 edge node in a K3s edge-cloud testbed, covering DNN inference with three vision models, multiple precisions, thread counts, and both local and serving scenarios. System-level MAPE is 6.3-9.4%, improving over CPU-stress and utilization-only baselines by 26.9% MAPE on average. We jointly report inference latency and energy per inference, exposing performance-energy trade-offs across workload configurations.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactiv...

Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al. · 0 citations
#edge computing Preprint Aug 2026

PRISM: Predictive Runtime In-place Scaling and Model Selection for Edge Microservices

PRISM, a prediction-guided runtime framework that jointly selects model variants and CPU allocations for containerized edge microservices, and adapts each pipeline stage in place and minimizes predicted CPU-package energy under deadline, resource, and offline model-level Quality of Result constraints is presented.

Uwe Gropengießer, Thomas Reuter, Dominik Schön et al. · 0 citations
#federated learning Open access Sep 2026

Optimizing resource allocation for federated LLM training via workload forecasting and autoscaling in edge environments

An integrated framework that combines LLM fine-tuning using federated adaptive local low-rank adaptation (FedALoRA) with forecasting-driven autoscaling policies is proposed, making it well-suited for federated-LLM Autoscaling in dynamic edge environments.

Bablu Kumar, Anshul Verma, Rajkummar Buyya · 0 citations
Book Open access Apr 2024

MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms

With its elastic power and a pay-as-you-go cost model, the deployment of deep learning inference services (DLISs) on serverless platforms is emerging as a prevalent trend. However, the varying resource requirements of different layers in DL models hinder resource utilization and increase costs, when DLISs are deployed...

Ji-Aang Duan, Shiyou Qian, Dingyu Yang et al. · 7 citations
#artificial intelligence Preprint Sep 2026

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

A carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware and matches cloud-level accuracy while reducing operational carbon emissions by $4\times on average.

Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, S. Tragoudas et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.