Preprint
Jul 2026
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives
Building on NCCL's device-side API, low-latency interfaces for constructing custom collective kernels are developed and used to implement new symmetric collectives in NCCL, demonstrating benefits for both AI inference and traditional HPC workloads.
Siyuan Shen, Anton Korzh, J. Bachan et al.
· 0 citations