Skip to content
Book Open access

CopyCat: Harvesting the Frequency Tax of Bulk Memory Copy

Sep 2026 · Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems · pp. 115-121 · 0 citations · 11 references

Abstract

Bulk memory copying (memcpy) is a dominant operation in modern data centers, driven by storage engines, in-memory databases, and caches. The emergence of heterogeneous memory (CXL, persistent memory, remote NUMA) further increases the volume of concurrent memory copy operations, making it easier to saturate the shared memory bandwidth. Past saturation, a higher CPU frequency no longer improves copy throughput. However, utilization-based OS control and Intel HWP both keep copying cores at the maximum frequency in the saturated copy phases that we evaluate, wasting up to 24% of CPU package energy on memory access stall cycles. We present CopyCat, a user-space runtime with a memcpy-like submission interface and an explicit completion operation that reclaims this wasted energy. It provides router selects between two complementary paths: a synchronous mode that down-clocks copying cores in place, and an asynchronous mode that offloads copies to dedicated low-frequency cores, freeing the issuing cores for compute. In a blob-cache server prototype, CopyCat saves 8% package energy in the asynchronous serve phase due to computing-memory accessing overlap that reduce the end-to-end execution time, while dedicated low-frequency cores further saves up to 22% under 1% throughput loss from frequency reduction.

Read PDF

Similar papers

Review Open access Aug 2026

Locality and Fast Paths in Buffer Pool Translation

Buffer pool translation—resolving an on-disk page identifier to an in-memory frame—was once a heavyweight operation whose bookkeeping consumed a substantial fraction of CPU cycles in early OLTP engines. Recent systems reduce this cost with techniques such as pointer swizzling, in-page hints, OS page-table mappings,...

Riki Otaki, Kathir Meyyappan, Aaron Elmore et al. · 0 citations
Review Aug 2026

Concurrency Response of Plain Global Loads on the NVIDIA H100

The main result concerns the plain-load path: attained LDG bandwidth peaks at a small offered per-thread load (K ~ 2) and then declines, by about 35% from K=2 to K=8 at the authors' primary configuration.

S. Manjunath, Rahul Ramachandra · 0 citations
Jul 2026

CrocSort: Resource-Efficient, Skew-Resilient Parallel External Merge Sort

CrocSort is presented, a byte-balanced parallel external merge sort with configurable memory and per-phase thread settings with practical resource-configuration rules for selecting these settings from input size, memory budget, and thread cap.

Riki Otaki, Charles Benello, Fuheng Zhao et al. · 0 citations
Preprint Sep 2026

Don't let your Memory defy you: Fragmentation-Aware Serverless Allocation with Elastic Memory Locality

Serverless platforms commonly rely on bundled, memory-centric configurations, where CPU capacity follows the specified memory size. Resource decoupling reduces this waste, but can create external fragmentation by producing diverse CPU-memory shapes that leave residual capacity stranded across nodes. Memory disaggregati...

Achilleas Tzenetopoulos, D. Masouros, Sotirios Xydis et al. · 0 citations
Preprint Sep 2026

Echo: Merging Host Device Buffers to Avoid Redundant Data Movement on Unified Memory SoCs

GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them before and after GPU execution. On UMA hardware these buffers reside in the same physical DRAM pool, so the copies consum...

Yu-Heng Zhu, Yan-Bo Zhao, Jia-Jia Li et al. · 0 citations
Preprint Aug 2026

MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC Workloads

MEPOWER is proposed, a flexible, model-based approach to exposing compute/data movement imbalance that characterizes the fine-grained memory behavior of parallel workloads that demonstrates a reduction in EDP on a range of HPC benchmarks with minimal impact on execution time when compared to the standard OS/hardware-ma...

Nanda Velugoti, Joseph Manzano, Andrés Márquez et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.