Wavel: A Fast and Efficient Compilation System for Wafer-Scale Accelerators
Abstract
Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a sharding-oriented view and therefore fail to fully leverage these emerging accelerators, leaving the dominant physical scheduling decisions unresolved. We make a key observation: at wafer scale, the central complexity of compilation shifts from finding a logical shard plan to constructing an executable schedule that fixes placement, layout alignment, and execution ordering. Based on this observation, we present Wavel, a fast and efficient compilation system for wafer-scale accelerators. Wavel introduces MeshIR to make executable schedules explicit, constructs them through two legal locality-preserving primitives, CUT and TUNE, and uses physical cost evaluation with structured pruning to keep search practical. On WSE-3 hardware, Wavel delivers 1.31×-1.35× higher mean per-setting throughput than the expert-optimized WaferLLM and up to 1.85× speedup over the Full-Chip baseline on dense models; calibrated cost-model results predict a 2.97× speedup on MoE workloads over the same baseline.