Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

Xema is presented, a memory-efficient diffusion serving system that exploits predictable tensor lifetimes for trace-guided memory optimization and introduces an offline planner that jointly selects parallelism, concurrency, and memory control under GPU memory and SLO constraints.

Xueze Kang, Guangyu Xiang, Suyi Li et al. · 1 citation
Preprint Aug 2026

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

Rollplex is presented, a runtime that decomposes the reference and training phase and moves the prefix computation into the rollout decode window and achieves speedup over serial colocation and disaggregation under the same GPU budget, while preserving the synchronous RL update.

Hanfeng Lu, Tianyu Feng, Suyi Li et al. · 0 citations