EcoxPU: Autonomous Multi-Agent Orchestration for Energy-Efficient Disaggregated LLM Inference on Heterogeneous xPUs
Large Language Model (LLM) inference is increasingly served on disaggregated clusters that combine different accelerator types, including NVIDIA GPUs, Huawei Ascend NPUs, and Kunlunxin AI processors. We show that these heterogeneous xPUs exhibit distinct energy-performance behaviors across inference phases and model mo...