LLM inference frequently exhausts GPU memory, forcing frameworks to stage data such as KV caches and intermediate tensors out of GPU HBM. Today this is done with GPUDirect Storage (GDS) over PCIe to NVMe SSDs, but NVMe bandwidth and latency remain a bottleneck for the fine-grained, high-frequency accesses these workloa...
Veerasenareddy Burru, Satananda Burla, Pradeep Kumar Nalla et al.· Proceedings of the 18th ACM...· 0 citations
Modern AI workloads demand microsecond-scale network reaction times, forcing data centers to offload congestion control algorithms (CCAs) to hardware. Simultaneously, emerging transport standards introduce diverse congestion signals like delay, CSIG, INT, and packet trimming. Understanding hardware-offloaded CCA adapta...
Meet Dadhania, R. K, Saptarshi Samanta et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.