Jul 2026· Proceedings of the International Conference on Parallel Processing· pp. 229-239· 0 citations· 36 references
Computer Science
TL;DR
A query density-driven key-range partitioning scheme that balances query density among PIM processors, allowing us to strike a balance between query load and data size via a parameter.
Abstract
Processing-in-Memory (PIM) systems, which consist of many processors with small local memory, have recently emerged as commercial products and attracted much attention as a means of overcoming the memory wall, particularly in the context of in-memory database technology. The state-of-the-art PIM-oriented index PIM-tree has been demonstrated to achieve asymptotically good spatiotemporal load balancing—query loads and data sizes are balanced among processors—for skewed queries, by trading spatial locality. Unfortunately, such a sacrifice of spatial locality hinders the PIM-oriented processing of range-aggregate queries. To achieve both spatiotemporal load balancing and efficiently executing range-aggregate queries on PIM systems, we develop a query density-driven key-range partitioning scheme. It balances query density among PIM processors, allowing us to strike a balance between query load and data size via a parameter. We then develop B\({}^\text{+}\)-Forest, a PIM-oriented B\({}^\text{+}\)-tree variant based on our partitioning scheme. Experimental results demonstrated that it exhibits higher skew resistance than a B\({}^\text{+}\)-tree based on space-constrained, query-load-balanced, density-unaware partitioning, and performance comparable to PIM-tree in point-get queries, as well as efficient support for range-aggregate queries.
This paper proposes an extension of query density-driven partitioning to support dynamically changing workloads and achieves significantly higher throughput than PIM-tree, a skew-resistant state-of-the-art data structure, during periods without workload changes.
A novel architecture called S !"#$, designed to enhance the performance of hash indexes in disaggregated memory, is introduced and the results show that S !"#$ outperforms state-of-the-art DM-optimized hash indexes by at most 6.7 → (RACE), 3.6 → (SepHash), and 1.8 → (Outback) in YCSB workloads, respectively.
Han-Tian Zha, Teng Ma, Bao-Tong Lu et al.· 0 citations
Staged query execution approach that interleaves query evaluation with fine-grained remote loading using random-access storage layouts and intermediate selection vectors to fetch only the necessary data for further processing is proposed, showing that fine-grained staged loading significantly reduces transferred data a...
David Loughlin, Holger Pirk· Datenbank-Spektrum· 0 citations
Workload Aware Column Imprint-Hash Join WACI-HJ is presented, which uses a workload-aware approach to accelerate hash joins and shows 1%, 38%, and 49% gain in CPU, RAM, and I/O, respectively.
With the rapid development of the big data industry, data volume across various industries has exploded, and large-scale datasets at PB and EB levels have become mainstream objects for data processing. Relying on core theories of distributed storage and query, this paper constructs an integrated collaborative optimizat...