BigMoMo: Efficient Inference of Large-Scale MoE with Speculative Decoding on Mobile Devices
Mixture-of-Experts (MoE) models expand language model capacity on smartphones, but expert offloading remains constrained by limited DRAM capacity and costly data movement. Sequential token routing couples expert execution to fragmented flash reads and multistage NPU preparation, leaving sparse computation stalled on we...