Sep 2026· Machine Learning and Knowledge Extraction· Vol 8, pp. 274· 0 citations· 44 references
TL;DR
Accelerated LLM, a modular architecture that replaces a single general-purpose LLM with an ensemble of task-specialised small language models (SLMs) governed by a neural query router and a Mamdani fuzzy inference system, achieves competitive or superior task-specific performance at a fraction of the parameter count.
Abstract
Large language models (LLMs) incur prohibitive computational costs when deployed as monolithic systems for multi-domain query processing. This paper proposes Accelerated LLM, a modular architecture that replaces a single general-purpose LLM with an ensemble of task-specialised small language models (SLMs) governed by a neural query router and a Mamdani fuzzy inference system. The router embeds each user query using a frozen sentence encoder and classifies it across four task domains—summarisation, translation, question answering, and text generation—routing confident queries directly to the corresponding SLM. Ambiguous queries are escalated to a three-input fuzzy logic system operating on Query Length, inter-Domain Overlap Score, and Classifier Confidence, enabling principled handling of imprecise inputs. A reinforcement-learning feedback loop, validated through a controlled pilot deployment, continuously refines the routing policy. The complete pipeline, including the sentence encoder, totals approximately 2.14 billion parameters—a 98.8% reduction relative to GPT-3.5 (175 B). The integration of fuzzy logic into the routing stage raises classification accuracy from 91.5% to 94.3% and reduces the hallucination rate to 9.8% (minor) and 6.4% (major). Evaluated on healthcare-augmented benchmarks against ChatGPT-3.5, Claude, Mistral 70B, and two contemporary compact models (GPT-4o-mini and Llama 3.1-8B-Instruct), Accelerated LLM achieves competitive or superior task-specific performance at a fraction of the parameter count. A small-scale pilot evaluation in the legal domain indicates that the routing and fuzzy logic components retain partial effectiveness beyond the primary healthcare setting, though full multi-domain validation remains future work.
This work compares Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Odds Ratio Preference Optimization (ORPO) using a novel reward modeling approach based on execution and semantic principles, revealing that while standard PPO suffers from reward sparsity and catastrophic collapse on 7B mod...
Noah Hampp, Katya Mirylenka, Michael R. Glass· Swiss Text Analytics Confere...· 1 citation
This work proposes ProRetrieval, which recasts the language model as a retrieval orchestrator: given a natural-language query, it synthesizes an executable program in a hybrid DSL interleaving SQL operators over structured fields with vector-retrieval primitives over text and images, with SQL itself providing the logic...
Chengcheng You, Zhen Sun, Yunhai Hu et al.· 0 citations
The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream i...
Zhi-Qiang Xia, Yang Li, Xin-Yuan Zhang et al.· 0 citations
The rapid adoption of Large Language Models (LLMs) for academic question answering has led to significant computational overhead, increased latency, and high energy consumption, even for queries that require minimal reasoning. Existing systems typically route all queries to LLMs without considering their complexity or...
Н. Н. Курьян, Devika Shibu, Libya Thomas et al.· International Conference Inn...· 0 citations
LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected, is proposed, suggesting that careful data curation, rather than scale, is the key to efficient Text-...
Hao-Yuan Ma, Heng-Wei Liu, Linjuan Wu et al.· 0 citations