Skip to content
Preprint

FlexPosit: Tunable Fractional Precision for LLM Inference Accelerators

Sep 2026 · 0 citations · 53 references
Computer Science

TL;DR

FlexPosit is a unified bit-serial systolic array with lightweight per-column decoders, unified Processing Elements (PEs), and a global precision controller, enabling tunable fractional precision while preserving fully regular systolic dataflow.

Abstract

Large language models (LLMs) offer remarkable capabilities but impose prohibitive compute and energy costs. Quantization governs the trade-offs between accuracy and hardware efficiency across granularity and bit-width. Finer granularity (e.g., group-wise) provides high accuracy but incurs scaling and control overhead, while coarser granularity (e.g., channel-wise) has lower overhead but loses accuracy at low precision. Meanwhile, mixed-precision quantization exposes rich accuracy-efficiency trade-offs algorithmically, but existing LLM accelerators remain limited to discrete precision modes, leaving the fractional design space between them unexplored. FlexPosit bridges these gaps through co-design of Posit-based quantization and a precision-tunable bit-serial architecture. Algorithmically, FlexPosit employs distribution-aware quantization with hardware-aligned, sensitivity-guided mixed-precision allocation, leveraging the Posit format's tapered precision to achieve group-wise-like accuracy with channel-wise-like regularity. Architecturally, FlexPosit is a unified bit-serial systolic array with lightweight per-column decoders, unified Processing Elements (PEs), and a global precision controller, enabling tunable fractional precision while preserving fully regular systolic dataflow. Across diverse LLMs, FlexPosit achieves near-FP16 accuracy with sub-5-bit fractional weights. It achieves 1.8x higher throughput and 1.2x lower energy than BitMoD (group-wise quantization), and 1.5x higher throughput and 2.0x lower energy than OliVe (channel-wise quantization), establishing a new Pareto frontier for precision-tunable LLM acceleration.

View source

Similar papers

Preprint Aug 2026

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference

AdaMX (Adaptive Microscaling), a heterogeneity-aware format and accelerator that removes 83% of the MXFP4 accuracy loss on commonsense and 82% on MMLU, and 43% and 27% of the NVFP4 loss across LLMs from 3B to 70B.

Junyi Luo, Xin Jiang, Tai-Hao Wen et al. · 0 citations
Preprint Aug 2026

FOCUS: FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling

FOCUS is proposed, a post-training quantization framework with end-to-end scale learning for FP4 Optimization via Coupled-Relaxation and Dual-Granularity Scaling, which relaxes the tight coupling between quantization and dequantization scales with a learnable full-precision coefficient, enabling more effective optimiza...

Xiang-Long Yan, Hong Liu, Cheng-Zhu Bao et al. · 1 citation
Book Open access Aug 2026

MECA-CiM: A Shared-MicroExponent-aware Configurable Analog Compute-in-Memory Macro for Efficient Inference

An analog CiM accelerator based on the SMX6 format, which extends the block floating-point representation with a lightweight microexponent shared by pairs of values is presented, demonstrating that micro-exponent-aware analog CiM with configurable granularity is an effective and practical design point for energy-effici...

Wonkyung Han, Dohyun Kim, Jihoon Park et al. · 0 citations
Preprint Sep 2026

Vortex: Bridging Extreme Compression and Efficient LLM Inference

This study addresses challenges with Vortex, an architecture compatible with systolic-array-based accelerators with minimal hardware overhead, bridging the gap between extreme compression and efficient inference, and proposes codebook-wise contextual sparsity to align with VQ execution.

Haoxuan Shan, Cong Guo, Bo-Wen Duan et al. · 0 citations
#machine learning Preprint Aug 2026

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

It is found that every one of five epilogue faults -- scale precision, double rounding, multiplication order, output truncation, fused ordering -- moves the output by at most a single bfloat16 spacing, and by exactly one whenever it moves it at all, across 5,880 cells.

Teng-Ruei Chen · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.