Jul 2026· Proceedings of the International Conference on Parallel Processing· pp. 229-239· 0 citations· 36 references
Computer Science
TL;DR
A query density-driven key-range partitioning scheme that balances query density among PIM processors, allowing us to strike a balance between query load and data size via a parameter.
Abstract
Processing-in-Memory (PIM) systems, which consist of many processors with small local memory, have recently emerged as commercial products and attracted much attention as a means of overcoming the memory wall, particularly in the context of in-memory database technology. The state-of-the-art PIM-oriented index PIM-tree has been demonstrated to achieve asymptotically good spatiotemporal load balancing—query loads and data sizes are balanced among processors—for skewed queries, by trading spatial locality. Unfortunately, such a sacrifice of spatial locality hinders the PIM-oriented processing of range-aggregate queries. To achieve both spatiotemporal load balancing and efficiently executing range-aggregate queries on PIM systems, we develop a query density-driven key-range partitioning scheme. It balances query density among PIM processors, allowing us to strike a balance between query load and data size via a parameter. We then develop B\({}^\text{+}\)-Forest, a PIM-oriented B\({}^\text{+}\)-tree variant based on our partitioning scheme. Experimental results demonstrated that it exhibits higher skew resistance than a B\({}^\text{+}\)-tree based on space-constrained, query-load-balanced, density-unaware partitioning, and performance comparable to PIM-tree in point-get queries, as well as efficient support for range-aggregate queries.
Near-memory acceleration, where a large number of compute nodes with limited memory process data in parallel, is a promising approach for in-memory databases. Therein, partitioning is required to balance data and queries for skewed workloads. However, existing load balancing methods lack efficient support for dynamic w...
A novel architecture called S !"#$, designed to enhance the performance of hash indexes in disaggregated memory, is introduced and the results show that S !"#$ outperforms state-of-the-art DM-optimized hash indexes by at most 6.7 → (RACE), 3.6 → (SepHash), and 1.8 → (Outback) in YCSB workloads, respectively.
Han-Tian Zha, Teng Ma, Bao-Tong Lu et al.· 0 citations
Disaggregated analytics systems separate compute, memory, and storage to improve elasticity and resource utilisation, making network transfer a central bottleneck. Existing systems typically access remote data at fixed coarse granularities, transferring entire files, columns, or column-chunks even when queries proces...
David Loughlin, Holger Pirk· Datenbank-Spektrum· 0 citations
Workload Aware Column Imprint-Hash Join WACI-HJ is presented, which uses a workload-aware approach to accelerate hash joins and shows 1%, 38%, and 49% gain in CPU, RAM, and I/O, respectively.
With the rapid development of the big data industry, data volume across various industries has exploded, and large-scale datasets at PB and EB levels have become mainstream objects for data processing. Relying on core theories of distributed storage and query, this paper constructs an integrated collaborative optimizat...