Skip to content
Preprint

Unifying In-Memory Data Analytics through Sparse Compilation

Sep 2026 · 0 citations · 47 references
Computer Science

TL;DR

This work introduces a novel intermediate representation (IR), grounded in relational algebra and sparse iteration theory, that provides a unified abstraction for data and computation that enables workload-agnostic, end-to-end optimizations across diverse analytics applications.

Abstract

As modern data analytics workloads become increasingly heterogeneous and hardware-intensive, achieving efficient multi-core performance across diverse applications remains an open challenge. We present Reffine, a compiler-based in-memory analytics engine that delivers high performance across a broad range of data analytics workloads. Reffine introduces a novel intermediate representation (IR), grounded in relational algebra and sparse iteration theory, that provides a unified abstraction for data and computation. This representation enables workload-agnostic, end-to-end optimizations such as operator fusion and automatic parallelization across diverse analytics applications. We further develop a sparse compiler backend that translates Reffine IR into hardware-efficient imperative code, achieving high multi-core performance without domain-specific implementations. On the TPC-H benchmark, Reffine outperforms the in-memory analytical database DuckDB by up to $24.9\times$ and the state-of-the-art compilation-based database Umbra by up to $3.2\times$. Reffine also achieves average speedups of $18.3\times$ and $47.9\times$ over Polars and NetworkX on streaming and graph analytics workloads, respectively. Source code: https://github.com/ampersand-projects/reffine

View source

Similar papers

Lessons Learned Building Cross-Architecture Analytical Engines

This thesis explores hardware-software co-design for data-intensive applications, targeting the unification of programming models using open standards and exploring experimental techniques for automated query synthesis, and presents X-BQSR, a holistic redesign of genomic base quality score recalibration pipelines.

I. D. Kabadzhov · 0 citations
Preprint Sep 2026

Decoupling Disaggregated Memory Optimizations from Indexing: A Compiler-Runtime Approach

Disaggregated memory (DM) decouples compute and memory into independently scalable pools, connected over a slower interconnect rather than a local bus. This decoupling is exactly what makes DM attractive--but it also means that every index must now reason explicitly about remote-memory access and its associated optimiz...

Xin-Peng Zhao, Ze-Ling Long, Chaichon Wongkham et al. · 0 citations
Book Open access Sep 2026

In-Copy Fusion: Runtime Argument Fusion for Efficient OpenMP GPU Offloading

In-Copy Fusion (ICF), a runtime optimization of OpenMP’s accelerator model, implemented in LLVM, that gathers eligible map clauses into staging buffers, orchestrating an optimal transfer pipeline through fusion without modifying application or kernel code, preserves OpenMP semantics with negligible overhead is presente...

Dionisis-Odysseas Sotiropoulos, Sara Royuela Alcázar, Eduardo Quiñones et al. · 0 citations
Preprint Aug 2026

A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

This work proposes FIBER, a new architecture that extends the GPU SIMT (single instruction, multiple thread) model, and extends the ISA, microarchitecture, and compiler to realize shared-register addressing, conflict-free operand delivery, and fiber-based program mapping.

Zihan Liu, Jingwen Leng, Yangjie Zhou et al. · 0 citations
Open access Sep 2026

Toward Disaggregated Analytical Database Systems in the AI Hardware Era

This paper identifies effective I/O-computation overlap as a key requirement for fully exploiting the AI data center stack, and outlines future research directions for next-generation analytical database architectures.

Ji-Gao Luo, Nils Boeschen, Muhammad El-Hindi et al. · 0 citations

Procyon: Promoting Fine-Grain Multi-Tenancy to Optimize Sparse Streaming Accelerators

Procyon, a fine-grain multi-tenancy framework that fuses the PE instruction streams of multiple workloads into a unified execution schedule, substantially reduces PE underutilization that results in 3 × speedup over state-of-the-art sparse streaming accelerators, and reaches a peak throughput of 61 .

Ubaid Bakhtiar, Jeremy Sha, Helya Hosseini et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.