Skip to content
Preprint

FABRICA: Agentic CUDA-to-CSL Translation and Optimization for Wafer-Scale Systems

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

These results identify base-model capability, target knowledge, execution feedback, and same-target measurement as central to cross-architecture kernel generation as central to cross-architecture kernel generation.

Abstract

Porting GPU kernels across architectures requires architectural remapping, not syntax substitution. CUDA encodes decomposition, locality, and synchronization through threads, blocks, and memory accesses; the Cerebras Software Language (CSL) requires explicit placement, distributed SRAM, fabric communication, event-driven tasks, and host/device contracts. We present FABRICA-Bench, 49 paired CUDA-to-CSL tasks, and FABRICA, an agentic framework combining target knowledge, execution, failure-directed repair, and correctness-gated optimization. On a fixed 28-task Level~1--3 core comparison with Claude Opus 4.8, FABRICA raises success from 6/28 to 26/28; 22 successful programs match or beat their CSL references. Across the 49-task coverage evaluation, 38 tasks produce a correct program; the final three tasks are evaluated over three seeds and pass 8/9 runs. For 27 generated/reference pairs with device-internal timing, geometric-mean speedup is 3.75$\times$ on the SDK simulator and 3.47$\times$ on WSE-3 hardware. With the executable workflow fixed, Claude Opus~4.8 passes 26/28 core tasks while the best open-weight model passes 2/28; retrieved Cerebras knowledge separately raises success from 1/15 to 7/15 on a Level~1--3 panel. These results identify base-model capability, target knowledge, execution feedback, and same-target measurement as central to cross-architecture kernel generation.

View source

Similar papers

Book Open access Sep 2026

Unlocking Software-defined GPU Fabric Scheduling in the LLM Era

Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We chara...

Dan-Yang Chen, Yu-Feng Gu, Yibo Huang et al. · 1 citation
Book Open access Sep 2026

Wavel: A Fast and Efficient Compilation System for Wafer-Scale Accelerators

Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...

Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al. · 0 citations
Preprint Aug 2026

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

CAKE, a compiler-agent co-design in which agents author CAKE IR, a typed, hardware-explicit schedule representation, exposes warp roles, memory movement, synchronization, and pipelines while supporting verification, cost modeling, and localized diagnostics.

Zihao Ye, Yingyi Huang, H. Jin et al. · 4 citations

Automated Generation of RISC-V Extensions with Formal Correctness Guarantees

Janus is presented, an LLM-assisted framework that synthe-sizes custom instructions integrated into the Ibex RISC-V core while keeping correctness outside the agent, demonstrating a practical path for using LLMs to explore ISA specialization without making the agent part of the trusted correctness boundary.

Elisavet Lydia Alvanaki, Jia-Kun Wang, Eugenio Muscinelli et al. · 0 citations
Preprint Sep 2026

Exo-GPU: Safe, Imperative, User-schedulable Programming for Tensor Cores

Exo-GPU, an imperative, low-level language that creates minimal abstraction over CUDA, is proposed, to treat parallelism and synchronization as mere annotations on sequential code rather than as fundamental control flow primitives, enabling verification that these constructs do not alter the program semantics.

David Akeley, Yuka Ikarashi, Jonathan Ragan-Kelley · 0 citations
#artificial intelligence Preprint Sep 2026

MaxKernel: Agentic Kernel Generation for TPUs

This work presents MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and a Graph-Based Autonomous...

Shang-Kun Wang, Nina Cai, Charles Hoong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.