Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Spexis: Speculative Lookahead Scheduling for LLM Inference

Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism, and uses lookahead scheduling to predict speculation quality and future memory pressure, to reduce wasted speculation, KV-cache eviction, and recomputation.

Hyungyu Jung, Jaehyeok Yu, Hoon Choi et al. · 0 citations
#machine learning Book Open access Apr 2026

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

DASH-Q is proposed, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares, which outperform other PTQ baselines in ultra low-bit regime and improves zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines.

Jaemin Kim, Sungkyun Kim, Junyeol Lee et al. · 0 citations
Open access Jul 2026

Runtime-Controlled Evaluation of Speculative Decoding Methods

Reported Speculative decoding (SD) speedups are difficult to compare because methods are commonly evaluated with different serving runtimes and configurations. We present SpecLLM, a pluggable evaluation framework that applies common scheduling, batching, KV-cache, CUDA Graph, and execution-backend policies across metho...

Sungkyun Kim, Jaemin Kim, Yeongpil Cho et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.