Decentralized Learning with Communication-Efficient Learned Gradient Sketches
Abstract
Decentralized learning eliminates single-point-of-failure risks inherent in parameter-server architectures, but its practical deployment is bottlenecked by the communication cost of exchanging full-precision gradients over peer-to-peer links. Existing compression techniques—random quantization, sparsification, and fixed low-rank projections—either sacrifice convergence quality under non-IID data or offer limited compression ratios. We propose Learned-Sketch Gradient Tracking (LS-GT), a framework that integrates a Micro-Block Diagonal sketching operator, learned online via a block-wise Sanger’s rule, into the gradient-tracking communication protocol. A Freeze-and-Update synchronization mechanism ensures all agents share an identical, periodically refreshed sketch, preserving the contraction properties required for error-feedback convergence. We prove that LS-GT converges linearly to a noise neighborhood of the global optimum at a rate governed by the learned subspace alignment factor, and that the amortized communication cost scales as \(\mathcal {O}(d/b)\) per iteration. Empirically, LS-GT matches uncompressed gradient tracking within 0.10 percentage points on ImageNet-1K at a 24x bandwidth reduction, and achieves near-lossless perplexity on Stack Overflow language modeling at 56x compression—outperforming QSGD, Top-K, and PowerGossip across all operating points.