Skip to content

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters

Jul 2026 · arXiv.org · Vol abs/2607.22432 · 1 citation · 56 references
Computer Science

TL;DR

This work presents TileSight, a tile-centric performance-modeling tool that elevates the tile from a programming primitive to an analysis primitive, and selects tile configurations competitive with strong vendor and expert baselines.

Abstract

Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernels. Performance-analysis tooling has not followed: programmers still rely on coarse roofline bounds, opaque ML predictors, or post-hoc profilers to understand kernel execution. This gap is acute for modern AI workloads, where kernel fusion and distributed inference depend on tensor cores, CUDA cores, cache hierarchies, memory pipelines, and inter-GPU networks. We present TileSight, a tile-centric performance-modeling tool that elevates the tile from a programming primitive to an analysis primitive. Within a GPU core, TileSight models compute-memory pipeline overlap; across cores, it models the cache hierarchy; across GPUs, it models inter-node communication. All layers share the tile abstraction: the intra-tile layer expresses work as a resource vector spanning network, memory, and compute pipelines; the inter-tile layer schedules dependent and ordered actions to expose legal overlap and infers multi-level cache hit rates from tile reuse distance; and the cross-device layer maps remote tensor accesses to placements and routes them through an alpha-beta stage cost. On A100, H200, B200, and B6000, TileSight predicts single-GPU kernel latency with 12.35% pooled mean absolute percentage error (MAPE), outperforming state-of-the-art baselines and transferring better across architectures. Its L2 cache-hit-rate predictions are within roughly one percentage point of measurements on every GPU. At up to 32 GPUs, TileSight achieves 16.18% weighted MAPE (wMAPE) on fused distributed kernels and 13.52% wMAPE on end-to-end vLLM serving. In optimization, TileSight selects tile configurations competitive with strong vendor and expert baselines. TileSight will be open-sourced upon publication.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granularity of kernels, on separate streams, reduces only part of this communication cost. Fused kernels often have better performance by transmitting each output tile as soon a...

Zi-Ming Mao, Yi-Han Zhang, S. W. Chew et al. · 0 citations
Preprint Aug 2026

A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

This work proposes FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model, and extends the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping.

Zihan Liu, Jingwen Leng, Yangjie Zhou et al. · 0 citations
Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations
Preprint Sep 2026

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems com...

Yao Fei, Gong-Ming Zhao, Hong-Li Xu et al. · 0 citations
Preprint Sep 2026

FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI Workloads

FlashGPU-sim is presented, an open-source, execution-driven, cycle-accurate GPU simulator for modern AI workloads that faithfully models modern hardware features such as asynchronous data movement, fine-grained synchronization, tensor-core execution, and distributed shared memory.

Si-Ying Yu, Yi-Xun Hong, Guo-Zhi Qiu et al. · 0 citations
Open access Aug 2026

Structure-Derived Bottleneck-Aware Scheduling for Multitasking MCM-GPUs

SA-Scheduler is presented, a structure-derived bottleneck-aware scheduling framework for multitasking MCM-GPUs that determines chip placement without hardware modification or runtime bottleneck profiling, and provides a principled and scalable foundation for multitasking on future MCM-GPUs.

Tie-Jian Zhang, Guangda Zhang, Lu Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.