Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· 0 citations· 67 references
Computer Science
TL;DR
Beeswax overlaps the execution of program stages, and thus hides cache miss overheads, and is applied to Katran, a load balancer deployed in production data centers, and BMC, a key-value store accelerator.
Abstract
eBPF has emerged as a popular platform for building high-performance I/O programs. However, eBPF's programming model limits the use of commonly used performance optimizations, leaving programs vulnerable to cache-miss-induced performance overheads. Specifically, the programming model restricts the choice of data structures used by a program, thus limiting the programmers' ability to improve data locality. Furthermore, it also makes it hard for programmers to overlap computation with data movement. In this paper, we present Beeswax an approach that addresses both of these restrictions. This approach requires programmers to adopt multi-phase data structures and partition programs into multiple stages. It overlaps the execution of program stages, and thus hides cache miss overheads. We designed the approach so that it could be applied to existing programs without requiring significant change. We have applied Beeswax to Katran, a load balancer deployed in production data centers, and BMC, a key-value store accelerator. We show throughput improvements of up to 99%.
A crash-consistent I/O cache is essential to ensure data integrity while optimizing performance. However, a general-purpose kernel fails to ensure integrity efficiently because of the high cost of synchronizing I/O requests with the many subsystems that rely on the cache. As a consequence, applications that must store...
Jana Toljaga, Nicolas Derumigny, Tara Aggoun et al.· Proceedings of the ACM SIGOP...· 0 citations
eBPF allows user-defined programs to safely extend Linux kernel functionality at runtime, but its final machine code comes from a compilation pipeline that differs from native targets, and how efficient that pipeline is has no clear reference point. Our work constructs one: using the standard LLVM x86 backend as an app...
Hoang Duong, Hao Sun, Zhen-Dong Su· Proceedings of the 4th Works...· 0 citations
This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.
Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al.· 1 citation
PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergenc...
Jaewoong Jung, Hyukyul Kang, J. Choi· Proceedings of the 18th ACM...· 0 citations
This work develops a novel low-overhead methodology for measuring the cost of GCs, by combining isolated thread monitoring using a combination of real hardware and calibrated cycle-accurate simulation, which allows for fine-grained analysis of modern GC overheads.
Sudhanshu Agarwal, Saugata Ghose· Proceedings of the ACM on Pr...· 0 citations
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.