K AIROX : Adaptive GPU–CPU Hybrid LLM Inference via Online Neuron Balancing
K AIROX introduces a Live Pipeline designed to prefetch neurons by predicting next-layer activation patterns, a mechanism that dynamically redistributes neurons between the GPU and CPU based on activation patterns, and a Temporal Activation Momentum cache policy to prioritize neurons with sustained utility while minimi...