Sep 2026· Proceedings of the 4th Workshop on eBPF and Kernel Extensions· 0 citations· 57 references
TL;DR
An eBPF-based scheduling framework is designed and implemented that complements existing kernel mechanisms to improve the accuracy of fine-grained CPU allocation while maintaining low overhead and demonstrates that the proposed approach significantly reduces resource overallocation and performance variability compared to conventional OS resource control mechanisms.
Abstract
The rise of cloud service models such as serverless computing and microservices has driven the adoption of fine-grained resource provisioning and billing, allowing users to request fractional CPU allocations at sub-core granularity. However, existing operating system mechanisms for CPU bandwidth control were designed for much coarser resource allocations. As a result, they can introduce substantial resource overallocation and performance variability when enforcing the small CPU shares required by modern cloud workloads. In this paper, we investigate whether eBPF can bridge this gap by enabling more precise and responsive CPU bandwidth enforcement. We design and implement an eBPF-based scheduling framework, μslice, that complements existing kernel mechanisms to improve the accuracy of fine-grained CPU allocation while maintaining low overhead. Our evaluation demonstrates that the proposed approach significantly reduces resource overallocation and performance variability compared to conventional OS resource control mechanisms.
This paper presents a novel hybrid segmented policy that reduces capacity requirements while keeping processing costs low and shows that dynamically adjusting the segment ratio in segmented policies based on historical workload patterns enhances efficiency.
The new EMC+ proposal is an OS‐driven elasticity manager for container‐based environments that continuously estimates idle core cycles left by regular (inelastic) applications, and reallocates idle cores to elastic ones, even during short time intervals, and has minimal impact on the performance and QoS of colocated in...
J. C. Saez, Carlos Bilbao, Manuel Prieto-Matías· Concurrency and Computation· 0 citations
In cloud data centers, the colocation of multiple services and applications on the same physical server is crucial for maximizing resource utilization and reducing utility costs. Unfortunately, contention for shared resources across cores, such as the Last-Level Cache (LLC), may lead to severe interference among applic...
Javier Aznal, J. C. Saez, Carlos Bilbao· Proceedings of the Internati...· 1 citation
OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.
Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al.· 0 citations
Large language models (LLMs) are increasingly deployed on edge nodes to support edge intelligence applications. To overcome limited GPU memory, offloading-based methods partition model parameters between the GPU and host memory, enabling inference on commodity hardware. However, deploying a single model instance using...
Zhen-Zheng Li, Zhiqing Tang, Jian-Xiong Guo et al.· IEEE Internet of Things Jour...· 0 citations
QMScaler is proposed, a QoS-driven microservice horizontal scaling framework based on Monte Carlo Tree Search (MCTS) that aims to minimize the number of container instances while meeting the QoS requirements of multiple application functions.
Tianyang Zheng, Pengfei Yang, Zhe Xu et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.