Skip to content
Preprint

How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation?

Aug 2026 · 0 citations · 23 references
Computer Science

TL;DR

This work uses Apache Spark running on Kubernetes as a case-study dataflow runtime and cluster resource manager to compare model-based energy estimates to Intel RAPL package and DRAM energy on an AWS bare-metal cloud and an on-premises cluster, comparing different CPU usage signals and memory coefficients.

Abstract

Distributed batch data processing applications are widely executed on cloud-based resources where restricted user access to node-level hardware energy counters hinders transparent sustainability accounting. Energy and carbon attribution methodologies therefore depend on power models and available resource utilisation traces, yet the accuracy of these estimates has to be validated while direct counters are available. In this work, we use Apache Spark running on Kubernetes as a case-study dataflow runtime and cluster resource manager to compare model-based energy estimates to Intel RAPL package and DRAM energy on an AWS bare-metal cloud and an on-premises cluster, comparing different CPU usage signals and memory coefficients. We show that external monitoring improves signed package-energy error relative to Spark task traces, reducing underestimation from -29.58% to -24.41% on AWS and from -24.00% to -16.22% on-premises.

View source

Similar papers

Preprint Sep 2026

Joule-Profiler: Profiling the Energy Consumption of Build Automation Tools Made Easy

Build pipelines are integral to modern software development, yet their energy footprint remains largely invisible to practitioners. Existing CI energy tools either rely on model-based estimation (due to hardware access restrictions in cloud runners) or report only total pipeline energy without decomposing it into meani...

Jérémy Woirhaye, François Gibier, Romain Rouvoy · 0 citations
2026

A Framework for Cloud Workload Allotment Using Real-Time Carbon Intensity of Energy Sources

Results show that the proposed framework for thermal-aware and carbon-efficient workload allocation could effectively select out the most suitable servers for workloads not only from viewpoint of carbon usage but also from the thermal perspectives, and at the same time, it also reduces the amount of cooling required an...

Chandan Hegde, Adarsh Bilimisi, Pruthvik J · 0 citations
Book Open access Sep 2026

Aethon: Performance-aware Memory Offloading for Co-running Applications in Public Clouds

Aethon is presented, a memory offloading system for public clouds that maximizes offloaded data while ensuring each application satisfies the service-level agreement (SLA).

Guang-Qiang Luan, Pu Pang, Quan Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

A carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture, combining edge and cloud LLMs on heterogeneous hardware and matches cloud-level accuracy while reducing operational carbon emissions by $4\times on average.

Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, S. Tragoudas et al. · 1 citation
Book Open access Sep 2026

Energy Efficiency in Actor Systems: A Comparative Study of Akka and Elixir

The energy consumption of digital infrastructures and applications constitutes a critical global concern. To build a sustainable digital future, software engineering must prioritize energy efficiency alongside traditional performance metrics. Actors provide an attractive model for distributed large-scale applications;...

I. Samus, Mario Südholt, C. De Roover et al. · 0 citations
Preprint Sep 2026

Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is Required

Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch schedulers such as Slurm. Growing demand for shared computational resources increases the need for interoperability between t...

M. Mačernis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.