The key insight in BOOST is to use kernel access patterns to make page allocation and runtime data management wave-aware, and applies modulo-based page placement that eliminates access-ratio variance.
A. Saxena, J. Ju, Hritvik Taneja et al.· 0 citations
PerfReasoning is introduced, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code, exposing the gap between plausible architectural reasoning and reliable performance-model construction.
Da Zhao, K. Sankaralingam, Christos Kozyrakis et al.· 0 citations
HiSparse is merged into upstream SGLang and evaluated across three sparse-attention families (DSA, NSA, and Quest) on H200, B200, and GH200 platforms: it improves peak generation throughput by up to 4.7x on long-context workloads while preserving comparable per-token latency and reducing time-to-first-token at high loa...
Valk, a performance analysis tool that combines data from multiple profilers, shows that when memory bandwidth is increased, kernels become compute bound, and makes three recommendations to fully utilize the GPUs' potential for relational workloads when the memory wall is removed.
S. Hepkema, Bo-Wen Wu, Christos Kozyrakis et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.