Towards High-Fidelity yet Low-Overhead Gradient Compression for In-Network Aggregation
As models continue to scale, distributed training remains bottlenecked by gradient synchronization overhead, even with In-Network Aggregation (INA). This has spurred extensive research on gradient compression, yet existing approaches fall short for INA: sparsification loses globally important gradients, uniform quantiz...