Skip to content

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Jul 2026 · arXiv.org · Vol abs/2607.13649 · 0 citations · 12 references
Computer Science

TL;DR

CIMERA is presented, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution.

Abstract

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to $25\times$ and $10\times$ higher energy efficiency for 1B and 13B models, respectively.

View source

Similar papers

Preprint Sep 2026

LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data...

Yimin Wang, Yue Jiet Chong, Xuan-Yao Fong · 0 citations
Book Open access Sep 2026

Analysis of shared memory between CPUs and GPUs

This work investigates the memory capabilities of the NVIDIA DGX Spark, a novel platform featuring a unified memory architecture where DDR memory is located on the CPU and is fully accessible from the GPU.

Silvia R. Alcaraz, S. Hepkema, Vasilis Mageirakos et al. · 0 citations
#edge computing Preprint Sep 2026

FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity

FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execu...

Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel · 0 citations
Preprint Aug 2026

MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC Workloads

MEPOWER is proposed, a flexible, model-based approach to exposing compute/data movement imbalance that characterizes the fine-grained memory behavior of parallel workloads that demonstrates a reduction in EDP on a range of HPC benchmarks with minimal impact on execution time when compared to the standard OS/hardware-ma...

Nanda Velugoti, Joseph Manzano, Andrés Márquez et al. · 0 citations
Review Aug 2026

AI Hardware Accelerators for Large Language Models: Architectures and the Memory Wall

Large language models (LLMs) place unprecedented and still-growing demands on the hardware that trains and serves them. This review surveys the full landscape of AI hardware accelerators for LLMs, including general-purpose GPUs, custom ASICs such as TPUs, Trainium, Groq, and Cerebras, reconfigurable FPGAs, processing-i...

Siddharth Patel, Rohit Singh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.