LogNexus: Effective Log Compression via Unified Redundancy Encoding
Abstract
State-of-the-art log compressors typically rely on a decoupled “parse-then-compress” workflow, where parsing is optimized for semantic accuracy (i.e., event identification) rather than storage efficiency. Through a comprehensive empirical study, we reveal that this architectural decoupling prevents the exploitation of deep correlations between static templates and dynamic variables. To address it, we propose LogNexus based on the principle of unified redundancy encoding, a new log compression paradigm that co-designs structural extraction and variable encoding. LogNexus constructs a Unified Redundancy Tree (URT) using a hierarchical strategy that progressively mines frequent “structure+variable” patterns in logs. Such a design captures deep contextual redundancies ignored by traditional methods while minimizing computational overhead by pre-emptively encoding dominant patterns. Extensive evaluation on 16 benchmark datasets demonstrates that LogNexus establishes a new state-of-the-art. It achieves the highest compression ratio on 14 datasets (outperforming baselines by 9.48%–89.13%) and the fastest speed (1.51×–40.06× faster than competitors). Furthermore, when configured in non-chunked mode to maximize global pattern discovery, LogNexus boosts its compression ratio by 285.13%, which is 27.08% higher than the best baseline, while retaining a 2.43× speed advantage. The decompression audit further shows that LogNexus successfully restores every token on all 16 datasets.