Skip to content
Book Open access

Analysis of shared memory between CPUs and GPUs

Sep 2026 · Proceedings of the 4th Workshop on Disruptive Memory Systems · pp. 52-55 · 0 citations · 14 references

TL;DR

This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.

Abstract

Heterogeneous architectures featuring CPUs and GPUs in one system are increasingly adopted for high-performance data processing, yet interconnect bandwidth and memory capacity remain primary bottlenecks on the GPU side. While high-end solutions like NVIDIA Grace Hopper mitigate these issues via specialized interconnects, their high cost limits widespread adoption. We investigate the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU. We analyze the performance and tuning of such a system and explore how data processing workloads can be best run on such shared memory architectures. The demonstration will showcase how memory is allocated, the performance implications of different configurations, the results of running a data analytics benchmark, and the tools used to run the benchmarks, insert instrumentation, and analyze the results.

Read PDF

Similar papers

Preprint Aug 2026

Over the Memory Wall, Into the Instruction Wall: The New Bottleneck in GPU Data Processing

Valk, a performance analysis tool that combines data from multiple profilers, shows that when memory bandwidth is increased, kernels become compute bound, and makes three recommendations to fully utilize the GPUs' potential for relational workloads when the memory wall is removed.

S. Hepkema, Bo-Wen Wu, Christos Kozyrakis et al. · 0 citations
Book Open access Aug 2026

A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU

This preliminary study evaluates coscheduling on the NVIDIA GH200 Superchip compared to a discrete H100 PCIe platform to suggest that integrated CPU–GPU platforms such as GH200 can improve both performance and programmability for coscheduled workloads.

Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah et al. · 0 citations
Conference Open access Sep 2026

Analyzing the Impact of Architectural Design Decisions on Performance Across Generations of NVIDIA GPUs

The high-performance computing industry is moving beyond an era in which each generation of GPU provides uniform performance gains across all applications. The growing importance of AI is driving GPU architecture towards greater specialization, with more silicon devoted to Tensor Cores and reduced-precision arithmetic....

Matthew Tindale, I. Karlin, Tobias Salamon et al. · 0 citations
Conference Sep 2026

MPTAC: Memory Controller Partitioning and Traffic-Aware Contention Control in GPUs

The advent of cloud-based artificial intelligence and the increased digitalization of embedded systems require powerful GPUs capable of simultaneously running kernels from different software providers. To accommodate the resource isolation and Execution Time Determinism (ETD) needed with the increasing number of kernel...

Vahid Geraeinejad, Paul Delestrac, Javier Barrera et al. · 0 citations
Open access Aug 2026

SAI: Virtualizing Shared Memory of GPU for AI workload acceleration

This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...

Hanqing Li, Tie-Jun Li, Sheng Ma et al. · 0 citations

FitFloat: Read/Write Random-Access Compressed Floating-Point Arrays for GPUs

FitFloat is presented, a drop-in floating-point array replacement supporting user-specified precision on GPUs with the goal of reducing storage requirements of scientific applications while maximizing performance over Unified Memory.

Andrew Rodriguez, Martin Burtscher · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.