Skip to content
Book Open access

GPU-Centric Stateless LLM Serving With GIGANETS

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · pp. 1988-1993 · 0 citations · 18 references
Computer Science

TL;DR

Giganetes is a GPU-centric architecture for stateless LLM serving that externalizes KV caches to a disaggregated remote memory pool via GPUDirect RDMA, eliminating session affinity constraints and enabling near-linear horizontal scaling in a Kubernetes-native deployment.

Abstract

Giganetes is a GPU-centric architecture for stateless LLM serving that externalizes KV caches to a disaggregated remote memory pool via GPUDirect RDMA. By treating remote memory as a GPU-addressable tier via GPUDirect RDMA, Giganetes eliminates session affinity constraints: any GPU can serve any request, enabling near-linear horizontal scaling in a Kubernetes-native deployment. A Scatter/Gather I/O interface bypasses the CPU and host memory entirely, achieving 52.4 GB/s application-level read throughput on our 4×200 Gbps RDMA testbed. A session-level metadata abstraction and proactive readahead mechanism reduce GPU bubbles by overlapping remote KV fetches with prefill computation and scheduling slack. On a 4-node H800 cluster, Giganetes delivers 33% higher throughput (QPS 2.4 vs. 1.8) and 1.75× lower P95 TPOT than PD-disaggregation with sticky sessions, with the gain driven by scheduling flexibility rather than faster transport alone.

Read PDF

Similar papers

Book Open access Sep 2026

Unlocking Software-defined GPU Fabric Scheduling in the LLM Era

Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We chara...

Dan-Yang Chen, Yu-Feng Gu, Yibo Huang et al. · 1 citation
Book Open access Sep 2026

AsymFlow: Enabling Long-Context LLM Serving via CPU-GPU Prefill-Decode Disaggregation

Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency...

Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al. · 0 citations
Book Open access Aug 2026

DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM Serving

DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...

Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.