Skip to content
Preprint

Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures

Sep 2026 · 0 citations · 58 references
Computer Science

TL;DR

Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures, is presented, suggesting that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.

Abstract

Modern AI and HPC accelerators increasingly expose dataflow features: software-visible mechanisms for data movement and overlap, such as inter-core communication through the on-chip network and intra-core asynchronous pipelining. These features shift scheduling responsibility from hardware to the compiler, and because placement, movement, and synchronization become software-visible, they also make the performance of static schedules predictable. Yet high performance on such hardware still relies on vendor-engineered kernel libraries or profile-based auto-tuning, whose embedded expert knowledge transfers poorly across architectures and algorithms. We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures. The central idea is to treat tile-based SPMD compilation as a hardware-explicit static optimization problem. Loom enumerates discrete spatial-mapping and communication candidates while keeping value parameters, such as tiling factors and pipeline knobs, symbolic within each candidate. From an explicit hardware description, it derives symbolic legality constraints and latency expressions, formulates one CP-SAT problem per schedule candidate, and jointly solves inter-core dataflow, intra-core asynchronous scheduling, and block sizes at compile time. On two Tenstorrent generations, Wormhole and Blackhole, Loom matches or exceeds the vendor-optimized TTNN library on GEMM, Flash Attention, and Flash Decode, out of the box and without per-shape profiling or profile-based platform-specific schedule tuning. These results suggest that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.

View source

Similar papers

Open access Aug 2026

Fine-Grained Structural Conflict Modeling for Compile-Time Instruction Scheduling on VLIW ASIPs

This paper proposes a compile-time instruction scheduling method that models sub-cycle resource usage and analyzes both data and structural dependencies at fine granularity, enabling precise detection of structural hazards in complex execution units.

Peng Hao, Shengbing Zhang, Xinbing Zhou et al. · 0 citations
Open access Aug 2026

Dolunay: Architectural Support for Independent Thread Scheduling in a RISC-V SIMT Accelerator

Dolunay is introduced, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator that employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms.

Ahmet Can, Erkan Uslu · 0 citations
Preprint Sep 2026

Compiler and Hardware Co-Design for Accelerator Architectures

Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardwar...

Karl Herman Krause, Emad Jacob Maroun, Martin Schoeberl · 0 citations
Nov 2026

RedPanda: A Unified Compilation Framework for Dataflow-Based CGRAs and High-Level Synthesis

Dataflow-based coarse-grained reconfigurable architectures (CGRAs) and dynamic high-level synthesis (DHLS) are both promising for accelerating applications with nontrivial control and memory behavior, but existing compilation flows are typically fragmented and often struggle with control handling, memory ordering, and...

Yi Huang, Xiang-Yu Kong, Jian-Feng Zhu et al. · 0 citations
Book Open access Sep 2026

Wavel: A Fast and Efficient Compilation System for Wafer-Scale Accelerators

Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...

Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al. · 0 citations
Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.