Skip to content
#small language model Open access

Accelerated LLM: A Fuzzy-Logic-Augmented Router Architecture for Efficient Multi-Domain Query Processing via Specialised Small Language Models

Sep 2026 · Machine Learning and Knowledge Extraction · Vol 8, pp. 274 · 0 citations · 44 references

TL;DR

Accelerated LLM, a modular architecture that replaces a single general-purpose LLM with an ensemble of task-specialised small language models (SLMs) governed by a neural query router and a Mamdani fuzzy inference system, achieves competitive or superior task-specific performance at a fraction of the parameter count.

Abstract

Large language models (LLMs) incur prohibitive computational costs when deployed as monolithic systems for multi-domain query processing. This paper proposes Accelerated LLM, a modular architecture that replaces a single general-purpose LLM with an ensemble of task-specialised small language models (SLMs) governed by a neural query router and a Mamdani fuzzy inference system. The router embeds each user query using a frozen sentence encoder and classifies it across four task domains—summarisation, translation, question answering, and text generation—routing confident queries directly to the corresponding SLM. Ambiguous queries are escalated to a three-input fuzzy logic system operating on Query Length, inter-Domain Overlap Score, and Classifier Confidence, enabling principled handling of imprecise inputs. A reinforcement-learning feedback loop, validated through a controlled pilot deployment, continuously refines the routing policy. The complete pipeline, including the sentence encoder, totals approximately 2.14 billion parameters—a 98.8% reduction relative to GPT-3.5 (175 B). The integration of fuzzy logic into the routing stage raises classification accuracy from 91.5% to 94.3% and reduces the hallucination rate to 9.8% (minor) and 6.4% (major). Evaluated on healthcare-augmented benchmarks against ChatGPT-3.5, Claude, Mistral 70B, and two contemporary compact models (GPT-4o-mini and Llama 3.1-8B-Instruct), Accelerated LLM achieves competitive or superior task-specific performance at a fraction of the parameter count. A small-scale pilot evaluation in the legal domain indicates that the routing and fuzzy logic components retain partial effectiveness beyond the primary healthcare setting, though full multi-domain validation remains future work.

Read PDF

Similar papers

2026

Optimizing Large Language Models for Robust Domain-Specific Text-to-SQL: From Prompting to Preference Alignment

This work compares Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Odds Ratio Preference Optimization (ORPO) using a novel reward modeling approach based on execution and semantic principles, revealing that while standard PPO suffers from reward sparsity and catastrophic collapse on 7B mod...

Noah Hampp, Katya Mirylenka, Michael R. Glass · 1 citation
Preprint Aug 2026

ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis

This work proposes ProRetrieval, which recasts the language model as a retrieval orchestrator: given a natural-language query, it synthesizes an executable program in a hybrid DSL interleaving SQL operators over structured fields with vector-retrieval primitives over text and images, with SQL itself providing the logic...

Chengcheng You, Zhen Sun, Yunhai Hu et al. · 0 citations
Preprint Sep 2026

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i...

Zhi-Qiang Xia, Yang Li, Xin-Yuan Zhang et al. · 0 citations
Conference Aug 2026

CHQR: A Context-Aware Hybrid Query Routing Framework for Energy-and Latency-Optimized Academic Question Answering

The rapid adoption of Large Language Models (LLMs) for academic question answering has led to significant computational overhead, increased latency, and high energy consumption, even for queries that require minimal reasoning. Existing systems typically route all queries to LLMs without considering their complexity or...

Н. Н. Курьян, Devika Shibu, Libya Thomas et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...

Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.