Skip to content

Author

Yixuan Tan

We have 2 of 18 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

DSpark is introduced, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification, and enables performance tiers that were previously unattainable, shifting the Pareto frontier of the DeepSeek-V4 serving system.

Xin Cheng, Xingkai Yu, Chenze Shao et al. · 24 citations · ⚡8
Book Open access Aug 2026

DualPath: Accelerating Agentic LLM Inference by Harvesting Disaggregated KV-Cache Storage I/O

DualPath is an inference system that breaks this bottleneck by introducing dual-path KV-Cache loading and enables a novel storage-to-decode path, in which the KV-Cache is loaded into decoding engines and then efficiently transferred to prefill engines via RDMA over the compute network.

Yongtong Wu, Shaoyuan Chen, Rilin Huang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.