Book
Open access
Aug 2026
OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving
Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism transitions, and live-migration-based load balancing to alleviate memory pressure and improve resource utilization.
Jingqi Feng, Guang Yang, Yukai Huang et al.
· Proceedings of the 32nd ACM... · 0 citations