Towards High-Fidelity yet Low-Overhead Gradient Compression for In-Network Aggregation
Abstract
As models continue to scale, distributed training remains bottlenecked by gradient synchronization overhead, even with In-Network Aggregation (INA). This has spurred extensive research on gradient compression, yet existing approaches fall short for INA: sparsification loses globally important gradients, uniform quantization distorts small values, and non-uniform quantization incurs prohibitive complexity. To address these issues, we propose SPIRE, a sketch-based gradient compression framework for INA with Persistent Importance Retention and Indices Encoding. Specifically, it preserves gradients of sustained importance across iterations, approximately stores gradient values through Prototype-Min Sketch, and losslessly compresses gradient indices via Block-Offset Encoding, thereby achieving high gradient fidelity with low compression overhead. We provide theoretical guarantees for SPIRE, including space complexity, error bounds for Prototype-Min Sketch, and lossless encoding proofs. The experimental results show that SPIRE improves iteration time by 34.49% and reduces compression error by 65.03% compared with existing solutions.