Preprint
Aug 2026
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding
EdgeXpert is proposed, a software-hardware co-designed LLM accelerator that resolves this incompatibility and achieves up to 56.3% latency reduction and 44.1% energy reduction compared to prior works, while maintaining near-baseline accuracy.
Sangwoo Ha, Hyunwoo Seo, Yurim Jo et al.
· 0 citations