Skip to content

Author

Tejinder Singh

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

The KV Cache Is the New Memory Wall

Autoregressive LLM inference at long context is bounded by memory bandwidth, not arithmetic throughput, and the binding resource shifts from model weights to the Key-Value (KV) cache as sequence length grows. For Llama-3-70B in BF16, the 140 GB weight footprint exceeds the 80 GB HBM of a single accelerator, and one 128...

Tejinder Singh · 0 citations
#machine learning Preprint Jul 2026

Optimizing AI Inference Across the Deployment Stack

A unified analytical treatment of inference optimization across the deployment stack with a three-layer taxonomy covering model-level techniques such as quantization, pruning, and distillation; compiler transformations such as graph fusion, layout optimization, and kernel autotuning; and system policies such as dynamic...

Tejinder Singh, John Pflueger, Jeebak Mitra et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.