Preprint
Aug 2026
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
AFlex is a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving and reduces energy per token by up to 49% over state-of-the-art disaggregated serving and 48% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.
Cun-Chen Hu, Liangliang Xu, Tianyu Liu et al.
· 0 citations