Coding agents powered by large language models (LLMs) repeatedly alternate between model inference and tool calls, creating long-lived sessions with reusable key-value (KV) states and asynchronous request resumptions. Logical readiness, however, does not ensure efficient admission in a shared serving system. Through di...
You-He Jiang, Fang-Cheng Fu, Bin-Hang Yuan et al.· 0 citations
Reducing LLM serving energy does not by itself guarantee lower deployment cost when electricity procurement exposes operators to unfavorable deviations from preset commitments. We study hourly commitments with positive, potentially asymmetric costs for overuse and underuse, and formulate energy-Performance-Aware Commit...
You Peng, You-He Jiang, Chen Wang et al.· 0 citations
GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents...
Dai-Feng Li, Huiqiang Jiang, Chengruidong Zhang et al.· 0 citations
As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: out...
Hao-Yu Zheng, Fang-Cheng Fu, Bin-Hang Yuan et al.· 0 citations
Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token probabilities and those assigned during rollout. Recomputing rollout log-probabilities with the policy-update backend avoi...
Ran Yan, You-He Jiang, Jia-Yi Nie et al.· 0 citations
AReaL-DTE is presented, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency and achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x within a cluster.
Yingqi Peng, Jia-Wei Zhang, Wenhao Zhou et al.· 1 citation
OpenTela is presented, a user-space orchestration overlay that turns existing fragmented HPC clusters into a unified, cross-institutional serving platform and provides a replicable blueprint for other sovereign AI initiatives to harness their own federated GPU infrastructure.
Xiaozhe Yao, Youhe Jiang, Ilia Badanin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.