Jul 2026· IEEE International Symposium on High-Performance Parallel Distributed Computing· pp. 609-610· 0 citations· 3 references
Computer Science
TL;DR
The challenges of scaling R-tree spatial search on a commercial Processing-in-Memory (PIM) system are studied, and parallel host-side aggregation, efficient result handling, and query-aware DPU assignment for scalable PIM-based spatial search are motivated.
Abstract
Spatial query processing is important in scientific, geospatial, and data-intensive applications. R-trees are widely used to index spatial objects, but their query-dependent traversal creates irregular work across different regions. This poster studies the challenges of scaling R-tree spatial search on a commercial Processing-in-Memory (PIM) system. Although PIM reduces CPU to memory data movement by executing search near memory, it does not remove full-pipeline overheads: the host still manages data placement, query batching, kernel launches, result retrieval, and aggregation. Our results show strong DPU-side search acceleration, with PIM kernel speedup ranging from about 20 × to 73 × , but end-to-end speedup is lower, ranging from 0.87 × to 11.29 ×. The runtime breakdown shows that CPU-side aggregation can dominate output-heavy workloads; on the Buildings dataset, aggregation accounts for 62.9% of total time, while DPU kernel time is only 4.4%. DPU-count scaling shows that more DPUs speed up the kernel, but end-to-end gains saturate due to full-pipeline overheads. We also observe a workload imbalance across the DPUs, with the ratio of maximum to mean hits reaching 29.1 × on Lakes. These findings motivate parallel host-side aggregation, efficient result handling, and query-aware DPU assignment for scalable PIM-based spatial search.
Algorithms for processing large-scale spatial datasets are of significant interest in both scientific research and industrial applications. The efficient implementation of such algorithms is crucial for modern data-intensive systems, and GPU-based parallel processing has emerged as a particularly effective approach for...
Ioannis Pateras, Polychronis Velentzas, M. Vassilakopoulos et al.· ISPRS International Journal...· 0 citations
A novel architecture called S !"#$, designed to enhance the performance of hash indexes in disaggregated memory, is introduced and the results show that S !"#$ outperforms state-of-the-art DM-optimized hash indexes by at most 6.7 → (RACE), 3.6 → (SepHash), and 1.8 → (Outback) in YCSB workloads, respectively.
Han-Tian Zha, Teng Ma, Bao-Tong Lu et al.· 0 citations
A novel code-generating engine with factorization that enables intra-query-parallelized query execution on factorized representations and generates code to overcome their CPU-unfriendly layout, offering a unified and scalable solution for modern workloads.
Stefan Lehner, Thomas Neumann· Proceedings of the VLDB Endo...· 0 citations
This paper proposes an extension of query density-driven partitioning to support dynamically changing workloads and achieves significantly higher throughput than PIM-tree, a skew-resistant state-of-the-art data structure, during periods without workload changes.
Staged query execution approach that interleaves query evaluation with fine-grained remote loading using random-access storage layouts and intermediate selection vectors to fetch only the necessary data for further processing is proposed, showing that fine-grained staged loading significantly reduces transferred data a...
David Loughlin, Holger Pirk· Datenbank-Spektrum· 0 citations
With the rapid development of the big data industry, data volume across various industries has exploded, and large-scale datasets at PB and EB levels have become mainstream objects for data processing. Relying on core theories of distributed storage and query, this paper constructs an integrated collaborative optimizat...