Skip to content
Open access

XK.jl: Composable and Portable Multi-GPU BLAS in Julia

Jul 2026 · JuliaCon Proceedings · Vol 8, pp. 206 · 0 citations · 25 references

TL;DR

This paper introduces XK.BLAS: the BLAS module of the Julia package XK.jl, which provides BLAS APIs that are portable across all three major GPU vendors (AMD, Intel, NVIDIA) for multi-device architectures for multi-device architectures.

Abstract

This paper introduces XK.BLAS: the BLAS module of the Julia package XK.jl. This module provides BLAS APIs that are portable across all three major GPU vendors (AMD, Intel, NVIDIA) for multi-device architectures. XK.BLAS is built on top of the XKRT tasking runtime systems and the XKBlas library. An XKBlas program is a sequence of task-generating routines with no explicit handling of memory (such as allocation and data movement) or scheduling decisions (such as target device selection). Memory is instead lazily allocated, migrated, and evicted, while tasks are automatically distributed across available devices. Tasks are composable, so that fine-grained dependencies are inferred from the memory regions they access, allowing concurrent data movement and kernel execution across different routines when the dataflow permits. We evaluate primitives of XK.BLAS and its use as a multi-GPU backend for the package Krylov.jl. We show up to 2.8 × speedup when scaling from 1 to 4 H100 GPUs. More importantly, XK.BLAS enables solving problems that would not otherwise fit on a single GPU, with no code changes at all.

Read PDF

Similar papers

Preprint Sep 2026

Interactive Debugger for Performance Portable Python HPC Kernels

We propose PKDB, the first interactive debugger for GPU and multithreaded low-level kernels written in Python. Python is widely used in high performance computing (HPC), with frameworks such as PyKokkos translating Python-embedded domain-specific languages to native code that runs across OpenMP-threaded CPUs and variou...

Ivan Grigorik, Gabriel Kosmacher, G. Biros et al. · 0 citations
#artificial intelligence Preprint Sep 2026

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granularity of kernels, on separate streams, reduces only part of this communication cost. Fused kernels often have better performance by transmitting each output tile as soon a...

Zi-Ming Mao, Yi-Han Zhang, S. W. Chew et al. · 0 citations
Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations
Preprint Aug 2026

GPU Offload in Rust: Portable, Safe, and Fast

This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.

Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al. · 1 citation
Preprint Aug 2026

GPU implementation of a resource-constrained virtual machine

This paper presents an OpenMP-style parallelism API for Uxntal, the stack-based assembly-style language for the Uxn platform, and demonstrates that exemplar code using this API can run at comparable performance even on an integrated GPU.

S. Li, Vladislav Brusokas, Andrei Ghita et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.