Skip to content

K AIROX : Adaptive GPU–CPU Hybrid LLM Inference via Online Neuron Balancing

· 0 citations · 60 references

TL;DR

K AIROX introduces a Live Pipeline designed to prefetch neurons by predicting next-layer activation patterns, a mechanism that dynamically redistributes neurons between the GPU and CPU based on activation patterns, and a Temporal Activation Momentum cache policy to prioritize neurons with sustained utility while minimizing transient, wasteful transfers.

View source

Similar papers

Oct 2026

SynergyScale: Optimizing Offloading and Task Partitioning for Efficient Model Training

Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...

Xiaoyang Sun, Jie Xu, Zheng Wang · 0 citations
Open access Sep 2026

ZenFlow: Enabling Stall-Free Offloading for LLM Training

Fine-tuning large language models (LLMs) often exceeds GPU memory limits, prompting systems to offload model states to CPU memory. However, existing offloaded training frameworks like ZeRO-Offload treat all parameters equally and update the full model on the CPU, causing severe GPU stalls, where fast, expensive GPUs si...

Ting-Feng Lan, Yu-Sen Wu, Bin Ma et al. · 0 citations
#machine learning Preprint Sep 2026

Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictab...

Sanjali Yadav, Bahar Asgari · 0 citations
#large language models Book Open access Aug 2026

NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching

NeuroPrefetcher is presented, a storage-backed LLM inference system that exploits that MLP activity during autoregressive decoding has strong temporal locality, and achieves 7.9-12.0x speedup over llama.cpp across constrained memory budgets.

Nobel Dhar, Md Romyull Islam, Xue-Chen Zhang et al. · 0 citations
Preprint Aug 2026

LazyTrain: Limited-resource Allocation toward Zero-waste Yield Optimization in Large Language Model Training

LazyTrain is proposed, an optimization layer over a layer-streaming executor that formulates checkpoint selection, activation placement, recomputation, and CPU-GPU-NVMe communication overlap as a mixed-integer scheduling problem, then executes the solved policy during training.

Xiao-Jun Wu, Ce-Hao Yang, Hong-Hao Liu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.