Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism, and uses lookahead scheduling to predict speculation quality and future memory pressure, to reduce wasted speculation, KV-cache eviction, and recomputation.
Hyungyu Jung, Jaehyeok Yu, Hoon Choi et al.· 0 citations
DASH-Q is proposed, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares, which outperform other PTQ baselines in ultra low-bit regime and improves zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines.
Jaemin Kim, Sungkyun Kim, Junyeol Lee et al.· EuroMLSys@EuroSys· 0 citations
Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across metho...
Sungkyun Kim, Jaemin Kim, Yeongpil Cho et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.