Skip to content

SymbolicLight V2: Hybrid Neuromorphic Architecture and Sparse Execution for Low-Energy Language Inference

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

The 194M-parameter model is implemented on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution to connect event sparsity to omitted computation and data movement in SymbolicLight V2.

Abstract

SymbolicLight V2 combines sparse event computation with continuous-state processing in a hybrid neuromorphic language architecture. Extending V1's spike-gated dual paths, it adds graded signed events at further projections and softmax-free local attention. We implement the 194M-parameter model on an Alveo U50C FPGA using digital fixed-point arithmetic and on an ARM CPU using sparse integer execution. Across three same-checkpoint FPGA implementations at 175 MHz, active-row weight gathering and valid-state KV loading raise decode throughput from 474.6 to 643.2 tokens/s for a 32-token prefix and 128 outputs. Estimated gross card energy falls from 0.06087 to 0.04407 J per generated token, a 27.6% reduction. Complete-request energy, including prefill, falls by 24.4-27.7% across three prefix lengths. An independent idle split attributes 82.8% of gross card energy to loaded idle, explaining the benefit of shorter token latency. Against the recorded RTX 5090 compiled-FP32 baseline, integer FPGA execution uses 89.1% less estimated card energy during short-context decode; arithmetic precisions differ, and the GPU baseline is not the lowest-energy tested configuration. On four Cortex-A76 cores of a ROCK 5T, complete requests reach 65.4 tokens/s at 9.80 W and 0.151 J per generated token at the adapter's AC input. These results connect event sparsity to omitted computation and data movement. The mechanisms also support other dedicated V2 implementations: increasing throughput by a greater factor than active power lowers energy per generated token. Evaluation holds the deployed checkpoint fixed; its quality trails a same-budget dense control, so the results do not establish equal-quality efficiency.

View source

Similar papers

#machine learning Preprint Aug 2026

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

This work introduces a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss, and positions sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.

Simon Richter, Ruhai Lin, Jason Yik et al. · 0 citations
Open access Oct 2026

Beyond binary SNNs: a hardware-algorithm co-design and architectural evaluation for graded spike networks

While Spiking Neural Networks offer a promising path toward energy-efficient edge intelligence, conventional Binary representations often suffer from high inference latency and discretization errors. Multi-level (graded) spike models address these limitations by increasing information density per pulse, yet they typi...

Wen-Fei Song, Andrea Castagnetti, Pierre-Emmanuel Novac et al. · 0 citations
#edge computing Preprint Sep 2026

FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity

FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execu...

Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel · 0 citations
#edge computing Review Open access Sep 2026

GPU and RISC-V acceleration for neuromorphic computing based on spiking neural networks: taxonomy, comparison, and open challenges

This survey reviews 71 studies published between 2015 and 2025 and organizes them into a taxonomy that classifies GPU works into simulation frameworks, training-acceleration techniques, large-scale and multi-GPU simulation, and application deployments, and RISC-V works into instruction-set extensions and tightly-couple...

Edris Zaman Farsa, Amirhossein Ilkhani, Marc Reichenbach et al. · 0 citations
Book Open access Sep 2026

Algorithm--Hardware Co-Design of Spiking Transformers for In-Memory Neuromorphic Edge Vision

The Temporal Hadamard Transformer is proposed, a spiking model co-designed with hardware for in-memory edge vision that integrates sparse algorithms with parallel hardware execution, providing a practical path toward real-time edge systems under realistic resource constraints and physical device non-idealities.

Yi-Xing Li, Wen-Hua Hu, Hao-Hui Peng et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.