Skip to content
Preprint

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

Abstract

Modern GPUs require not only SIMT-style parallelism but also software-managed concurrency between compute and data movement to reach maximum performance. Performance engineers must reason about subdividing work into the hierarchy of computation resources (threads, warps, warpgroups, blocks, clusters), and, in many cases, also must use asynchronous tensor core and memcpy instructions on different levels of the memory hierarchy (registers, tensor core accumulators, shared memory, global memory). Unlike CPUs, where out-of-order execution is managed by hardware and hidden from programmers, GPUs expose explicit instruction reordering to software through these asynchronous instructions. Well-established GPU programming languages generally offer either direct low-level control without safety guarantees (e.g., CUDA C++ inline assembly or intrinsics) or easier-to-analyze, high-level abstractions (e.g., Triton's tile-based model) that hide asynchronous instructions in the compiler backend, which may prevent performance engineers from maximizing performance by tuning critical details. We propose Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA. Our key idea is to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics. The benefit is twofold: programmers can reason about code without hidden control flow or mutation, while allowing the Exo-GPU compiler to verify sequential-parallel equivalence--guaranteeing that parallel execution is functionally equivalent to its sequential interpretation. We used Exo-GPU to author GEMM kernels for the H100 GPU, using wgmma, TMA, and split-k. Our kernels achieved over 80% of theoretical peak on large problem sizes, in some cases outperforming the vendor-provided CUBLAS library.

View source

Similar papers

Preprint Aug 2026

A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

This work proposes FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model, and extends the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping.

Zihan Liu, Jingwen Leng, Yangjie Zhou et al. · 0 citations
Preprint Aug 2026

GPU Offload in Rust: Portable, Safe, and Fast

This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.

Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al. · 1 citation
Open access Aug 2026

Dolunay: Architectural Support for Independent Thread Scheduling in a RISC-V SIMT Accelerator

Single-Instruction Multiple-Thread (SIMT) architectures have revolutionized data-parallel computing by providing a high-throughput abstraction that simplifies vector management. However, traditional stack-based SIMT models do not support intra-warp synchronization primitives such as mutexes and spin-locks. This work in...

Ahmet Can, Erkan Uslu · 0 citations
Book Open access Sep 2026

PLINK: A GPU-Initiated I/O Platform Exploiting NVMe Parallelism

As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps onl...

Jaewoong Jung, Hyukyul Kang, J. Choi · 0 citations
Book Open access Sep 2026

Computation Is Fast, Use Threadlet!: Efficient Threading for μs-Scale Computing via OS/Hardware Co-Design

The conventional threading model multiplexes software threads onto hardware cores. This model inherently suffers from the overhead of 1) context switching, 2) scheduling, and 3) system event notifications (e.g., I/O interrupts). As computing enters μs-scale, such overheads become the key bottleneck of a wide range of d...

Yi-Ming Yao, Xiao-He Qin, Yi Fan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.