Oct 2026· Proceedings of the ACM on Programming Languages· Vol 10, pp. 735 - 764· 0 citations· 104 references
TL;DR
This work develops a novel low-overhead methodology for measuring the cost of GCs, by combining isolated thread monitoring using a combination of real hardware and calibrated cycle-accurate simulation, which allows for fine-grained analysis of modern GC overheads.
Abstract
Garbage collection is an essential part of modern managed languages, such as Java, that are used in billions of devices and in a large variety of settings. Multiple garbage collectors (GCs) have been developed over the last several decades, in an attempt to optimize across a complex design space that includes memory footprint for GC metadata, thread concurrency with the mutator (i.e., the application), performance, and the footprint of stale data. While aggregate runtime metrics have been used to guide modern GC design, it has been difficult to use such metrics to capture the fine-grained performance and energy impact that GC execution has on the memory system. We develop a novel low-overhead methodology for measuring the cost of GCs, by combining isolated thread monitoring using a combination of real hardware and calibrated cycle-accurate simulation, which allows us to perform fine-grained analysis of modern GC overheads. We use our methodology to make several observations about the overheads of six modern Java GCs on the memory system, including: (1) the GCs introduce a substantially higher overhead on L3 cache accesses compared to L1 cache accesses at all levels of GC pressure analyzed; (2) GC loads require more time on average to be serviced, a cost that increases with reduction in GC pressure; (3) GC accesses generally do not improve application cache hits, as the potential benefits of prefetching application data are counteracted by GC–application interference; and (4) modern highly-concurrent GCs account for most of the useless prefetching due to L2 hardware prefetching triggered by workload execution. Our work highlights the assumed memory system overheads of GC not captured by existing metrics, and aims to encourage future work in optimizing GCs for the memory system.
The transition of Java applications from monolithic, bare-metal enterprise deployments to cloud-native, containerized environments (e.g., Kubernetes) has introduced severe memory management challenges. Traditional Java Virtual Machine (JVM) heuristics—designed for high-throughput, multi-gigabyte monolithic runtimes—oft...
Tulsiram Sharma, Shubham Agrawal· International Journal of Res...· 0 citations
LMTracer is presented, a fine-grained and real-time performance profiling framework for production LLM services that embed profiling logic into the execution through graph-embedded probing and streaming buffered profiling data to CPUs on demand to keep the execution of user kernels uninterrupted.
Wei Liu, Yong-Chao He, Bo-Han Zhao et al.· Proceedings of the ACM SIGOP...· 0 citations
Beeswax overlaps the execution of program stages, and thus hides cache miss overheads, and is applied to Katran, a load balancer deployed in production data centers, and BMC, a key-value store accelerator.
Farbod Shahinfar, Marco Molè, Aurojit Panda et al.· Conference on Applications,...· 0 citations
It is found that although the new protection measures aimed at stopping popular heap exploit techniques or restricting the exploit strategy space can effectively defend against the exploitation of most vulnerabilities, the attack strategies proposed by this work can still make these vulnerabilities exploitable again.
Torchy is presented, a tracing JIT compiler for PyTorch, one of the mainstream eager-mode frameworks, that achieves similar performance as data-flow frameworks, while providing the same semantics of straight-away execution.
eBPF allows user-defined programs to safely extend Linux kernel functionality at runtime, but its final machine code comes from a compilation pipeline that differs from native targets, and how efficient that pipeline is has no clear reference point. Our work constructs one: using the standard LLVM x86 backend as an app...
Hoang Duong, Hao Sun, Zhen-Dong Su· Proceedings of the 4th Works...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.