Skip to content
Preprint

QEmbed: A Deep Learning Based Cardinality Estimator for Efficient Query Processing

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

This paper proposes a deep learning model formally called QEmbed, built upon the Masked Autoencoder for Distribution Estimation (MADE) auto-regressive framework to learn joint data distributions for selectivity estimation, and designs a hybrid encoding scheme that combines one-hot and embedding encodings.

Abstract

Cardinality estimation is at the core of any commercial database system for efficient query processing. Over the decades, non-learning-based estimation techniques (e.g., histogram-based, sampling-based) have been widely used in both commercial and open-source database platforms. However, these techniques are only effective when the number of columns in a table is small, as they cannot properly capture dependencies between multiple attributes. Recently, learning-based approaches have been shown to perform significantly better than the heuristic methods that have been used for the past three decades. Despite this success, existing learned models often struggle to balance memory efficiency and accuracy when dealing with datasets that mix high and low cardinality attributes. In this paper, we propose a deep learning model formally called QEmbed. Our model is built upon the Masked Autoencoder for Distribution Estimation (MADE) auto-regressive framework to learn joint data distributions for selectivity estimation. To improve data representation and overcome the limitations of using a single encoding method, we design a hybrid encoding scheme that combines one-hot and embedding encodings. This hybrid design enables QEmbed to retain fine-grained attribute information for smaller domains while capturing compact semantic patterns for large, sparse domains. We capture attribute correlations by factoring the joint data distribution into a series of conditional probabilities. This approach naturally accommodates both point and range queries. Through extensive experiments, we show that while QEmbed faces a latency trade-off on extremely wide schemas, it provides highly reliable cardinality estimates overall. A key advantage of our model is that it reduces extreme tail errors (maximum Q-errors), avoiding catastrophic estimation failures on complex, highly correlated workloads.

View source

Similar papers

#edge computing Preprint Aug 2026

MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection

This paper presents MetaSieve, a metapath selection layer that determines which metapaths to retain and which to prune, and shows that MetaSieve consistently reduces per-epoch training time by large margins while maintaining and often improving accuracy.

Fahim Shahriar Khan, Ashraf Aboulnaga · 0 citations
#artificial intelligence Preprint Sep 2026

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

A lightweight GLiClass-based router is introduced, a lightweight GLiClass-based router that assigns a suitability score to each inference-time model label without autoregressive generation, and a released 0.6B-parameter checkpoint combines a Qwen3 decoder with a shallow bidirectional scorer.

Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko et al. · 0 citations
Open access Aug 2026

Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation

REP-LIE leverages the gradients of LoRA low-rank matrices to estimate the importance of weights without requiring full gradient computation, and a stability score is introduced, serving as the basis for iterative pruning of unimportant model parameters.

Peng Liu, Hui-Bing Zeng, Yi-Qun Zhang et al. · 0 citations
Preprint Aug 2026

Every Expert Counts: ExactMoE for Memory-Efficient W4A16 Inference

ExactMoE, an inference design that applies symmetric group-128 four-bit weight quantization only to routed experts, stores those experts in kernel-native MARLIN form in pinned host memory, and executes all selected experts through a configurable GPU-resident slot cache and fused grouped MoE kernels, identifies a practi...

Amjad Saab · 0 citations
#artificial intelligence Preprint Sep 2026

MAxBench: A Multinomial Concept Recovery Benchmark

This work introduces MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation, and finds that affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces.

Divya Appapogu, Freya Behrens, Yonatan Belinkov et al. · 0 citations
Preprint Sep 2026

MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference

Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to...

Ting-Hao Wang, Yi-Chen Guo, Qi-Zhe Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.