These results identify base-model capability, target knowledge, execution feedback, and same-target measurement as central to cross-architecture kernel generation as central to cross-architecture kernel generation.
Yuebo Luo, E. Huerta, Venkat Vishwanath et al.· 0 citations
Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy. For the same SpMM on the same matrix, cuSPARSE exhibits a 350x performance gap between CSR and Blocked-ELL. Our study of mult...
HLSmith, an expert-guided framework for translating C/C++ programs into optimized HLS accelerators, is presented and evaluated on PolyBench against ChatHLS, a leading prior agent-orchestration framework for HLS accelerator development.
Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.