Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration
Xema is presented, a memory-efficient diffusion serving system that exploits predictable tensor lifetimes for trace-guided memory optimization and introduces an offline planner that jointly selects parallelism, concurrency, and memory control under GPU memory and SLO constraints.