Skip to content
Preprint

Tree-Structured Vector Quantization For Efficient And Progressive Image Compression

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

Tree-VQ is proposed, a progressive tree-structured vector quantization framework for learned image compression that codes progressive continuation decisions and routed branch refinements using only causally available decoded contexts and achieves a superior performance--efficiency trade-off.

Abstract

Vector-quantization based image compression has achieved strong rate--distortion performance, yet most of them still produce a separate compressed representation for each target bitrate. Such variable-rate behavior allows one model to operate at multiple rates, but it does not necessarily provide a progressive bitstream whose prefixes are themselves decodable and can be refined by appending additional bits. We propose \textbf{Tree-VQ}, a progressive tree-structured vector quantization framework for learned image compression. Tree-VQ organizes discrete codewords as a hierarchical binary tree and represents each latent token by a routed root-to-leaf path. Crucially, every prefix of this path corresponds to a valid quantized representation, so shallow internal nodes serve as coarse reconstruction codes and deeper nodes provide successive refinements. This allows a compressed image to be decoded from an early prefix and progressively improved as more branch symbols are received, rather than being re-encoded for different target rates. To make this structure practical for compression, we introduce a prefix-compatible tree entropy model that codes progressive continuation decisions and routed branch refinements using only causally available decoded contexts. We further use rate-aware refinement scheduling to decide which spatial blocks should receive additional tree bits under a given prefix budget, and hierarchical prefix supervision to ensure that internal nodes are directly decodable at low rates. Experiments show that Tree-VQ achieves a superior performance--efficiency trade-off, delivering the best perceptual compression results with much fewer parameters and lower latency than competing methods.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

XOR-Trellis: Ultra-Low-Complexity Dequantization and Curvature-Aware Hadamard-Free LLM Quantization

Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient pa...

Xiao-Fan Que, Nir Elkayam, Spandan Pyakurel et al. · 0 citations
Preprint Sep 2026

VQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGA

Learned image compression (LIC) is hard to deploy on severely resource-constrained FPGAs, since how fast it actually runs depends not just on arithmetic count, but also on memory traffic, imbalance between different operations, and how the hardware batches its work. We present VQ-LIC, an asymmetric edge-cloud codec in...

Muhammad Fahd Ibrahim Bhatti, Abdullah Bin Faisal, Ahsan Usman et al. · 0 citations
Review Aug 2026

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

This work identifies and formalizes the principle that organizes the Great Inversion, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening.

Ehsan Jokar · 0 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

This work introduces Matryoshka Hash Representations (MHR), a two-stage procedure that separates full-width training from prefix organization and strengthens two common pipelines: shortlisting candidates for full-precision reranking, and pruning a low-storage graph index such as LEANN.

Pei-Chun Hua, Yun-Ming Xiao · 1 citation
#artificial intelligence Preprint Sep 2026

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

This work builds a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning.

Hao-Ran Jin, Kang-Qi Zhang, Ji-Rong Yang et al. · 0 citations
#machine learning Preprint Sep 2026

When Does Low-Bit Quantization Preserve the Decisions of Vector Search?

Low-bit quantization can achieve high recall on some vector representations and fail sharply on others, while average distortion and global rank correlation do not explain the difference. We study quantized vector search at the level of the comparisons consumed by ranking and graph-pruning algorithms. Our first result...

Wen-Xuan Xiao, Xu Cao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.