Book
Open access
Aug 2026
Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation
This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.
Ying Wan, Yuchen Xu, Chuwen Zhang et al.
· Proceedings of the ACM SIGCO... · 0 citations