Preprint
Jul 2026
SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs
Evaluation across multiple datasets shows that SelectInfer achieves significant reductions in memory footprint and computation while preserving task performance, making it a practical step towards enabling LLM deployment on edge devices.
Huzaifa Shaaban Kabakibo, Eric Schniedermeyer, Artem Burchanow et al.
· 0 citations