Sep 2026· Proceedings of the 4th Workshop on Disruptive Memory Systems· pp. 52-55· 0 citations· 14 references
TL;DR
This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.
Abstract
Heterogeneous architectures featuring CPUs and GPUs in one system are increasingly adopted for high-performance data processing, yet interconnect bandwidth and memory capacity remain primary bottlenecks on the GPU side. While high-end solutions like NVIDIA Grace Hopper mitigate these issues via specialized interconnects, their high cost limits widespread adoption. We investigate the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU. We analyze the performance and tuning of such a system and explore how data processing workloads can be best run on such shared memory architectures. The demonstration will showcase how memory is allocated, the performance implications of different configurations, the results of running a data analytics benchmark, and the tools used to run the benchmarks, insert instrumentation, and analyze the results.
Valk, a performance analysis tool that combines data from multiple profilers, shows that when memory bandwidth is increased, kernels become compute bound, and makes three recommendations to fully utilize the GPUs' potential for relational workloads when the memory wall is removed.
S. Hepkema, Bo-Wen Wu, Christos Kozyrakis et al.· 0 citations
This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.
Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah et al.· Workshop Proceedings of the...· 0 citations
The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....
Matthew Tindale, I. Karlin, Tobias Salamon et al.· Inquiry@Queen's Undergraduat...· 0 citations
The advent of cloud-based artificial intelligence and the increased digitalization of embedded systems require powerful GPUs capable of simultaneously running kernels from different software providers. To accommodate the resource isolation and Execution Time Determinism (ETD) needed with the increasing number of kernel...
Vahid Geraeinejad, Paul Delestrac, Javier Barrera et al.· IEEE International Conferenc...· 0 citations
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
FitFloat is presented, a drop-in floating-point array replacement supporting user-specified precision on GPUs with the goal of reducing storage requirements of scientific applications while maximizing performance over Unified Memory.
Andrew Rodriguez, Martin Burtscher· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.