Skip to content

An Introduction to Compression-Based Machine Learning

Sep 2026 · 0 citations · 65 references
Computer Science

TL;DR

This work introduces and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware.

Abstract

Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly circular dependence has unrealized potential in modern artificial intelligence and machine learning, and we survey and formalize the various strategies that have been used to leverage compression for machine learning. We introduce and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware. We find that varying these design choices yields accuracy gains of up to 0.62.

View source

Similar papers

Conference Open access Aug 2026

Mixture-of-Experts-Based Entropy Model for Learned Image Compression

Learned image compression has seen significant progress in recent years with the development of end-to-end learned models that achieve better compression efficiency than state-of-the-art conventional methods. Recently, Mixture of Experts (MoE) approaches have seen promising results in NLP and computer vision tasks. In...

Jonas Brenig, R. Timofte · 1 citation
#artificial intelligence Preprint Sep 2026

Entropy-Punctured Bloom Filters for Memory-Efficient Machine Learning

Memory-efficient feature representations are increasingly important in machine learning settings where storage, transmission cost, bandwidth, or privacy constraints limit access to raw data. Bloom Filter (BF) encodings provide compact probabilistic representations of engineered features, but their behavior under struct...

J. Cartmell, M. Cardei, Ionut Cardei · 0 citations
Open access Oct 2026

LogNexus: Effective Log Compression via Unified Redundancy Encoding

State-of-the-art log compressors typically rely on a decoupled “parse-then-compress” workflow, where parsing is optimized for semantic accuracy (i.e., event identification) rather than storage efficiency. Through a comprehensive empirical study, we reveal that this architectural decoupling prevents the exploitation of...

Yang Liu, Kai-Ming Zhang, Zhuang-Bin Chen et al. · 0 citations
#natural language process... Preprint Sep 2026

FlexComp: One Model for Every Ratio in Context Compression

Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose...

Kai-Yan Zhao, Zhong-Tao Miao, Akiko Aizawa et al. · 1 citation
Preprint Aug 2026

Lossy Compression via Sparse Regression Codes: Generalized Construction and Finite-length Bounds

This work generalizes the SPARC construction, and considers the class of \emph{additive orthogonal} regression codes, of which standard SPARCs are a special case, and derives nonasymptotic bounds on the squared-error distortion by tracking the evolution of the encoding residual across stages.

G. Reeves, R. Venkataramanan · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.