Skip to content
Book Open access

Low-bit and Sparsified Gradient Communication for Accelerating Distributed Deep Learning with Convergence Guarantees

Sep 2026 · Proceedings of the International Conference on Parallel Processing · pp. 1041-1050 · 0 citations · 33 references

TL;DR

A low-bit and sparsified gradient communication algorithm with convergence guarantees, called QTopKA2A, which enables effectively overlapping communications with both feed-forward and backpropagation computations through a decomposed communication framework.

Abstract

Communication poses a dominant bottleneck in distributed data parallel training with synchronous stochastic gradient descent, whereas the traditional AllReduce collective used for gradient synchronization limits the efficient utilization of communication compression strategies. In this paper, we propose a low-bit and sparsified gradient communication algorithm with convergence guarantees, called QTopKA2A, which enables effectively overlapping communications with both feed-forward and backpropagation computations through a decomposed communication framework. Specifically, top-k sparsified gradients are communicated via AlltoAll in the backward phase, while low-bit quantized gradients are exchanged via AllGather in the forward phase. Residual-based error feedback mechanisms are employed in both sparsification and quantization stages to compensate for compression errors and guarantee convergence with significantly higher compression ratios. In addition, we leverage an effective tensor fusion to reduce communication startup overhead. Experimental results demonstrate that QTopKA2A preserves dense-level accuracy while achieving up to 4.74 × end-to-end training speedup over existing dense and compressed baselines on a 32-GPU cluster.

Read PDF

Similar papers

Book Open access Aug 2026

Aquavit: Ascending Quantization for Communication-Efficient Vast-Scale Distributed Training

Training Large Foundation Models (LFMs), including Large Language Models and Vision-Language Models, on massive distributed GPU clusters is increasingly bottlenecked by communication overhead. While frameworks like ZeRO++ employ static quantization to reduce communication volume, they suffer from a rigid trade-off: agg...

Hong Huang, Jia-Xun Ye, Jin-Hai Yang et al. · 0 citations
Book Open access Sep 2026

Decentralized Learning with Communication-Efficient Learned Gradient Sketches

Decentralized learning eliminates single-point-of-failure risks inherent in parameter-server architectures, but its practical deployment is bottlenecked by the communication cost of exchanging full-precision gradients over peer-to-peer links. Existing compression techniques—random quantization, sparsification, and fixe...

Ze-Hua Cheng, Wei Dai, Jia-Hao Sun · 0 citations
Book Open access Sep 2026

SSQT: A Hardware-Friendly Fusion Compression Framework of Structured Sparsification and Sensitivity-Driven Quantization for Large-Scale Language Models

Large language models (LLMs) are often memory-bandwidth bound during autoregressive decoding, so reducing weight storage does not automatically produce parallel speedup when the compressed representation is irregular. We present SSQT, a post-training framework that jointly applies hardware-aligned structured sparsifica...

Qian-Sheng Song, Guo-Lin Tang · 0 citations
Book Open access Sep 2026

Towards High-Fidelity yet Low-Overhead Gradient Compression for In-Network Aggregation

As models continue to scale, distributed training remains bottlenecked by gradient synchronization overhead, even with In-Network Aggregation (INA). This has spurred extensive research on gradient compression, yet existing approaches fall short for INA: sparsification loses globally important gradients, uniform quantiz...

Jin Wang, Chen Zhu, Jin-Bin Hu · 0 citations
Book Open access Aug 2026

Towards Network-Efficient Cross-Regional Inference via Learned Activation Compression

Feather is presented, a system that reduces this overhead by compressing intermediate activations before transmission and reconstructing them before downstream layers resume execution, consistently outperforming existing compression schemes, including linear autoencoder, PCA, and SVD.

Regan McDonald, M. Rego, Ertza Warraich et al. · 0 citations
Preprint Aug 2026

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) are typically applied in isolation, and each hits an accuracy-efficiency wall. This thesis proposes the"Compression Trini...

Mohammad Mozaffari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.