Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.
Abstract
Modern GPUs require not only SIMT-style parallelism but also software-managed concurrency between compute and data movement to reach maximum performance. Performance engineers must reason about subdividing work into the hierarchy of computation resources (threads, warps, warpgroups, blocks, clusters), and, in many cases, also must use asynchronous tensor core and memcpy instructions on different levels of the memory hierarchy (registers, tensor core accumulators, shared memory, global memory). Unlike CPUs, where out-of-order execution is managed by hardware and hidden from programmers, GPUs expose explicit instruction reordering to software through these asynchronous instructions. Well-established GPU programming languages generally offer either direct low-level control without safety guarantees (e.g., CUDA C++ inline assembly or intrinsics) or easier-to-analyze, high-level abstractions (e.g., Triton's tile-based model) that hide asynchronous instructions in the compiler backend, which may prevent performance engineers from maximizing performance by tuning critical details. We propose Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA. Our key idea is to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics. The benefit is twofold: programmers can reason about code without hidden control flow or mutation, while allowing the Exo-GPU compiler to verify sequential-parallel equivalence--guaranteeing that parallel execution is functionally equivalent to its sequential interpretation. We used Exo-GPU to author GEMM kernels for the H100 GPU, using wgmma, TMA, and split-k. Our kernels achieved over 80% of theoretical peak on large problem sizes, in some cases outperforming the vendor-provided CUBLAS library.
This work proposes FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model, and extends the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping.
Zihan Liu, Jingwen Leng, Yangjie Zhou et al.· 0 citations
This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.
Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al.· 1 citation
Single-Instruction Multiple-Thread (SIMT) architectures have revolutionized data-parallel computing by providing a high-throughput abstraction that simplifies vector management. However, traditional stack-based SIMT models do not support intra-warp synchronization primitives such as mutexes and spin-locks. This work in...
Ahmet Can, Erkan Uslu· WiPiEC Journal - Works in Pr...· 0 citations
As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps onl...
Jaewoong Jung, Hyukyul Kang, J. Choi· Proceedings of the 18th ACM...· 0 citations
The conventional threading model multiplexes software threads onto hardware cores. This model inherently suffers from the overhead of 1) context switching, 2) scheduling, and 3) system event notifications (e.g., I/O interrupts). As computing enters μs-scale, such overheads become the key bottleneck of a wide range of d...
Yi-Ming Yao, Xiao-He Qin, Yi Fan et al.· Proceedings of the ACM SIGOP...· 0 citations