Spexis: Speculative Lookahead Scheduling for LLM Inference
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism, and uses lookahead scheduling to predict speculation quality and future memory pressure, to reduce wasted speculation, KV-cache eviction, and recomputation.