Skip to content
Open access

The Hidden Cost of Container Networking: CNI Overlay Throughput Constraints in Kubernetes-Managed Apache Spark Versus Native Hadoop MapReduce on Bare-Metal Edge Hardware

2026 · IEEE Access · Vol 14, pp. 134602-134611 · 0 citations · 26 references

Abstract

This paper examines the impact of migrating large-scale data processing frameworks from bare-metal servers to cloud-native environments on system performance. Specifically, we compare traditional Hadoop MapReduce deployed on bare-metal infrastructure with Apache Spark running on Kubernetes, using a 200 GB TeraSort workload. Our analysis shows that Spark reduces overall execution time by approximately $4.5\times $ . Through examination of system telemetry across CPU, memory, disk, and network components, we identify the hardware and software factors contributing to this performance improvement. The findings indicate that Hadoop’s legacy architecture causes prolonged CPU I/O wait states due to extensive persistent disk writes. In contrast, Spark uses DDR5 RAM and the operating system’s page cache to reduce reliance on physical storage, sustaining near-continuous user-space compute. However, this in-memory processing introduces an overhead from the Kubernetes Container Network Interface (CNI). While Spark eliminates the physical storage bottleneck, its software-defined networking overlay limits intra-cluster data shuffling, resulting in network throughput that is substantially lower than native bare-metal speeds. We conclude that although cloud-native Spark offers a clear performance advantage over traditional MapReduce for speed-sensitive tasks, achieving high throughput requires high-bandwidth memory provisioning and the deployment of kernel-level eBPF networking techniques to mitigate Kubernetes network overhead.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.