This paper identifies effective I/O-computation overlap as a key requirement for fully exploiting the AI data center stack, and outlines future research directions for next-generation analytical database architectures.
Abstract
The AI hardware boom has driven modern data centers toward HPC-style architectures centered on GPU clusters, RDMA-capable networks, and high-throughput NVMe storage. While designed primarily for training and inference, this infrastructure also creates new opportunities for designing the next generation of scalable database systems for analytical workloads on top of such data centers. In particular, the combination of GPU-centric computation, fast networking, and fast storage enables disaggregated architectures that extend beyond single-node, GPU-memory-resident execution.
This paper discusses the challenges and design considerations of analytical query processing on such disaggregated GPU-centric systems. We examine how modern networks and storage enable distributed execution and out-of-memory processing, and how their interaction shapes end-to-end performance. Our recent results show that naïve use of existing I/O abstractions can underutilize both compute and I/O bandwidth due to insufficient overlap between computation and data movement. We therefore identify effective I/O-computation overlap as a key requirement for fully exploiting the AI data center stack, and outline future research directions for next-generation analytical database architectures.
The growing disparity between processor core scaling and memory bandwidth has exposed the physical and economic limits of processor-centric database architectures. While Compute Express Link (CXL) and other emerging technologies enable a necessary shift toward memory-centric, disaggregated topologies, it also introdu...
Yi Jiang, Hamish Nicholson, Anastasia Ailamaki· Datenbank-Spektrum· 0 citations
CAISA is introduced, a composable AI systems architecture that enables disaggregated memory expansion for large-scale AI workloads using CXL-based shared memory and coupling memory isolation with software-managed data orchestration decouples compute and memory resources while preserving efficient data movement, providi...
Divya Kiran Kadiyala, Lianjie Cao, Jinsun Yoo et al.· Conference on Applications,...· 0 citations
This work introduces a novel intermediate representation (IR), grounded in relational algebra and sparse iteration theory, that provides a unified abstraction for data and computation that enables workload-agnostic, end-to-end optimizations across diverse analytics applications.
This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.
Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al.· Proceedings of the 4th Works...· 0 citations
A memory-centric design model in which a database expresses what it needs from memory as declarative properties and a runtime resolves them against whichever fabric is present, and three architectural principles from recent work that instantiate this approach at different layers of the stack are consolidated.
SolidAttention is introduced, an LLM inference engine which addresses limitations through a tight co-design of dynamic attention sparsity algorithms and SSD-based storage management and minimizes SSD-induced blocking latency.
Xin Zheng, Dong-Liang Wei, Jian-Xiang Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.