Preprint
Aug 2026
Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference
ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practical memory-transfer-throughput frontier for complete-expert MoE inference.
Amjad Saab
· 0 citations