How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation?
This work uses Apache Spark running on Kubernetes as a case-study dataflow runtime and cluster resource manager to compare model-based energy estimates to Intel RAPL package and DRAM energy on an AWS bare-metal cloud and an on-premises cluster, comparing different CPU usage signals and memory coefficients.