Skip to content

Author

Sai-Qian Zhang

We have 3 of 21 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference

Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption. Prefill--decode (PD) disaggregation has emerged as a prevalent architecture for large-scale inference serving. However, determining the...

Ming-Yuan Yan, Hai-Yu Wang, Lin-Xuan Biao et al. · 0 citations
#machine learning Preprint Sep 2026

Acceptance-Aware Draft Model Training for Speculative Decoding

Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are verified by the target model in a single forward pass. Its speedup is largely determined by the acceptance length, yet existing draft-model training methods mainly opti...

Tian-Hua Xia, M. Ganesan, Yi-Fei Feng et al. · 0 citations

Stealing Split Learning Bottom Models by Recovering Embedding Geometry

VENOM first learns a contrastive space over server-observed embeddings, then builds a neighborhood graph and trains a surrogate bottom model to match targets and respect local geometry via a neighbor-matching loss alongside pointwise and feature-shape alignment to preserve the relational structure that defenses fail to...

Qinbo Zhang, Yanhang Shi, Ziying Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.