Skip to content

Accelerating String-Heavy Queries with LLM Token Tables

Jul 2026 · Proceedings of the VLDB Endowment · 0 citations · 66 references

Abstract

Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which require full-block decompression before processing. In this paper, we explore the potential of repurposing Large Language Model (LLM) tokenizers as a lightweight string compression scheme for databases, similar to FSST, but with a global token table shared across all tables and columns. Operators such as joins and aggregations can exploit this consistent encoding to defer decompression and process encoded values directly. We implement a global token table based on GPT-4's tokenizer in Umbra and demonstrate execution time improvements of up to 2× on string-heavy workloads, while reducing storage and memory consumption by up to 1.65×. Tokenizers integrate well with other compression algorithms, such as FSST, OnPair, or Zstd, while maintaining good compression ratios and high decompression throughput exceeding 6 GB/s on a single CPU core.

View source

Similar papers

Preprint Aug 2026

Direct-Operable SIMD Bit-Slicing: A Framework for Memory-Efficient Predicate Evaluation

A novel framework that utilizes the Project Panama Vector API to perform predicate evaluation directly over bit-sliced, compressed data streams by transposing standard row-oriented data into parallel bit-planes to demonstrate a mechanism to evaluate complex filters using SIMD instructions without requiring prior decomp...

A. Mathiyazhagan · 0 citations
Open access Sep 2026

G-CasDec: General Cascaded Decompression on GPUs

GPU-accelerated analytical query processing is often limited by both GPU device memory capacity and host-to-device data transfer time. Modern data compression techniques, such as cascaded lightweight compression, can mitigate these issues. However, existing designs all exhibit critical tradeoffs on compression ratios,...

Yong-Qi Zhuo, Xin-Yu Zeng, Huan-Chen Zhang et al. · 0 citations

S !"#$ : A Scalable and Resize-optimized Hash Index on Disaggregated Memory

A novel architecture called S !"#$, designed to enhance the performance of hash indexes in disaggregated memory, is introduced and the results show that S !"#$ outperforms state-of-the-art DM-optimized hash indexes by at most 6.7 → (RACE), 3.6 → (SepHash), and 1.8 → (Outback) in YCSB workloads, respectively.

Han-Tian Zha, Teng Ma, Bao-Tong Lu et al. · 0 citations
Preprint Aug 2026

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

DataKernelBench is introduced, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair and finds that higher-performing implementations commonly use kernel fusion and execut...

Gokul Karthik Kumar, Yotam Perlitz, Corey Lammie et al. · 0 citations
Jul 2026

The Data World is Not Flat: Efficient Factorized Execution for Relational Systems

A novel code-generating engine with factorization that enables intra-query-parallelized query execution on factorized representations and generates code to overcome their CPU-unfriendly layout, offering a unified and scalable solution for modern workloads.

Stefan Lehner, Thomas Neumann · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.