Skip to content

Category

small language model

813 papers

#small language model Preprint Aug 2026

FOCUS&RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models

A token-level analysis of this failure mode is presented by viewing decoding as a dynamical process that enters and persists in a small set of recurrent contexts and shows that persistence is controlled by the escape mass assigned to plausible alternatives within the token sampling set.

Junyoung Lee, Se-Hee Park, Shinhyoung Jang et al. · 0 citations
#small language model Preprint Aug 2026

Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

This work plants a controllable latent variable inside natural-looking text and arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.

Alexandru-Iulius Jerpelea · 0 citations
#small language model Open access Aug 2026

Intent Drift in LLM-Assisted Brain Computer Interface Communication: An In-Silico Benchmark Under Simulated Decoder Corruption

Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed, and this does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed.

A. Gorenshtein, M. Omar, E. Jia et al. · 0 citations
#small language model Review Open access Aug 2026

Artificial intelligence-enabled sustainability in sme supply chains: a systematic and bibliometric literature review

The findings indicate that machine learning is the dominant AI technology in SME supply chains, primarily used for forecasting, inventory management, process monitoring, logistics optimization, anomaly detection, and operational decision support, and economic and environmental sustainability dimensions receive substantially greater attention than social sustainability.

L. Fonseca, Luca Esposito, T. Murino et al. · 0 citations

Low Carbon Scheduling of Integrated Energy System Based on Large Language Model-Embedded Multi-Agent Reinforcement Learning

The complex multi-energy coupling characteristics inherent to integrated energy system (IES) present unprecedented challenges for the implementation of low-carbon scheduling. Existing optimization methods often exhibit limitations in system scalability, algorithm adaptivity, and carbon reduction efficacy for complex IES. This paper proposes a Large Language Model (LLM)-Embedded Multi-Agent Reinforcement Learning (LEMARL) to address the aforementioned issues. The proposed method integrates the global perception capability of LLMs with the dynamic optimization capability of MARL. Specifically, the LLM-Embedded module generates high-quality reward functions and policy frameworks from a global perspective, while the MARL module leverages these LLM-generated strategies for distributed interactive iterations—greatly enhancing computation efficiency and scalability. Simulation results demonstrate that LEMARL reduces carbon emissions by 7.76% and simultaneously decreases operating costs by 4.49% in a small-scale IES. Furthermore, LEMARL also exhibits superior applicability and scalability in large-scale IES of the IEEE 141-bus power grid integrated with 51-node thermal system.

Chen Xia, Tong Gou, Yinliang Xu et al. · 1 citation

DuoPIM: RRAM–DRAM Hybrid PIM Acceleration for Flexible-Batch LLM Decoding

Transformer-based large language models (LLMs) primarily consist of weight-intensive fully connected (FC) layers and cache-dependent attention layers. While batching significantly enhances the throughput of FC layers, it paradoxically increases the cache demands of attention layers. This provides no performance benefit and creates substantial memory pressure. Consequently, existing graphics processing unit (GPU)-based LLM acceleration systems face throughput limitations from batch size constraints. Even when DRAM-based processing-in-memory (PIM) is employed to accelerate attention, the utilization remains extremely low under small batch sizes, which is unsuitable for low-batch scenarios. Fortunately, the emerging nonvolatile resistive random access memory (RRAM) technology offers batch size-insensitive acceleration for FC layers through highly parallel in situ computations by eliminating weight loading overhead. This insight leads us to propose a hybrid approach: RRAM for FC layers and DRAM PIM for attention layers to overcome batch size limitations. However, merely scaling existing RRAM architectures misaligned with LLMs’ computation and storage demands will result in prohibitive overheads. Meanwhile, existing DRAM-based PIMs suffer from poor resource utilization due to the computational pattern of attention layers. Implementing an effective scheduling strategy is equally crucial to harness the potential of the hybrid PIM system. To address these challenges, we present DuoPIM, a novel RRAM–DRAM hybrid PIM architecture optimized for LLM decoding. We introduce novel architectural innovations for both the RRAM and DRAM PIM components to address the challenges posed by LLMs. Specifically, we decouple RRAM’s storage and computing capabilities within a hierarchical architecture, implement minimal modifications to DRAM PIM to support online softmax, and devise dedicated strategies across multiple architectural levels to enhance overall resource utilization. Evaluations demonstrate DuoPIM’s ability to fully leverage computing capacity across various batch sizes.

Xiaotian Sun, Xinyu Wang, Wanqian Li et al. · 0 citations
#small language model Preprint Aug 2026

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.

Leonardo Duart, T. Fonseca, T. Chacon · 0 citations
#small language model Preprint Aug 2026

Localize-Then-Decide Guarantees for LLM Judgments

This work proposes a Localize-Then-Decide framework, which restores the monotonic relationship between confidence and disagreement risk and enables high-probability agreement guarantees in large language models.

Xinyue Li, Yi Zhou, Guanqun Cao et al. · 0 citations
#small language model Preprint Aug 2026

Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

The constrained case of this"model size vs. inference compute"trade-off, in which the model outputs are constrained by a strict grammar at inference time, is examined, which demonstrates that the constrained trade-off behaves differently from the unconstrained trade-off.

Ty Chermsirivatana, John MacCormick · 0 citations
#small language model Preprint Aug 2026

RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models

RTLGuard leverages a teacher-student framework designed to sanitize compromised RTL generation models by fine-tuning a small-scale,"clean"teacher model on a limited set of trusted RTL data, and incorporating feature alignment and knowledge distillation to suppress malicious behaviors.

Mahshid Rezakhani, K. Azar, H. Kamali · 0 citations
#small language model Preprint Aug 2026

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

Spectral-Aware Muon is introduced, which holds the head at the Muon scale and amplifies the bulk using a static spectral prior, and both variants outperform tuned AdamW and Muon (Scion implementation) baselines in all evaluated model-scale and batch-size configurations.

Xiaodong Wu, Wenyi Yu, Chao Zhang et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.