Optimizing AI Inference Across the Deployment Stack
A unified analytical treatment of inference optimization across the deployment stack with a three-layer taxonomy covering model-level techniques such as quantization, pruning, and distillation; compiler transformations such as graph fusion, layout optimization, and kernel autotuning; and system policies such as dynamic...