Skip to content

Category

small language model

813 papers

#small language model Preprint Aug 2026

Disentangling Optimization Scale from Preference Scale in DPO

This work shows that $\beta$ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size, and proposes a centered-softplus reformulation that is argmin-equivalent to DPO for $\beta>0$, while making the inverse preference-noise-scale and learning-rate effects explicit and independently tunable.

Ivan Kruzhilov · 0 citations
#small language model Preprint Aug 2026

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.

Yu-Fan Wu, Yinghui He, Zhengyi Hu et al. · 1 citation
#small language model Preprint Aug 2026

Reservoir: A Large-Scale Simulated Dataset for Training and Evaluating Epidemiological Models

Large-scale, standardized datasets have driven many advances in AI-based scientific modeling, from protein structure prediction to natural language processing. Infectious disease epidemiology is increasingly adopting AI methods for forecasting, surveillance, and outbreak analytics, but the time-series data available to train them remains orders of magnitude smaller than the corpora behind the advances seen in other fields. Because the scope of real-world epidemiological data cannot practically reach the scale needed to train truly large-scale AI methods, simulated data provides a possible alternative. Here we introduce Reservoir, a large open simulator and dataset of realistic epidemic simulations in which every trajectory carries complete ground-truth labels, including quantities that cannot be measured directly in a real outbreak, such as true infection counts, time-varying reproduction numbers, and counterfactual intervention effects. Reservoir is generated by a stochastic simulator with realistic noise and reporting artifacts, together with interventions with configurable timing, compliance, and age-dependent efficacy. The current release contains 500,000 outbreak trajectories spanning one billion simulated days across diverse pathogen characteristics, population structures, and intervention regimes. Reservoir enables counterfactual experiments, surveillance-design studies, and training of epidemic models at a scale real-world datasets cannot provide.

Carson Dudley, Reiden Magdaleno, Marisa Eisenberg · 0 citations
#small language model Preprint Aug 2026

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

This work introduces Visual Retrieval Heads (VRHs), a small subset of attention heads that are causally responsible for grounding text descriptions to image regions, and shows that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads.

Chanho Park, Daehyeon Choi, Jihyun Lee et al. · 0 citations
#natural language process... Preprint Aug 2026

Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

This report presents an open pretraining recipe that trains a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs, and derives a Puro Cost Scaling Law that relates training cost to average model performance.

Kairong Luo, Jia-Rui Cui, Yao-Rui Yin et al. · 0 citations
#small language model Open access Aug 2026

The Performance of Large Language Models in Extracting Intestinal Symptoms From Electronic Health Records: Retrospective Observational Study

This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.

Xinyue Zhang, Quanyu Wang, Beibei Liu et al. · 0 citations
#small language model Preprint Aug 2026

Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning

CMPM, a Chinese Multi-Panel Meme benchmark with 1,214 annotated samples covering five structural types, ordering dependency, panel-order constraints, and optional comment context, is introduced and results indicate that canonical-display accuracy is not by itself evidence of order understanding.

Hai-Han Li, Hai-Hao Li, Zhengjie Xu et al. · 0 citations
#small language model Preprint Aug 2026

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

The effect of introducing the manager-worker scaffold over a shared filesystem workspace, with no training and no per-benchmark tuning, measured against the same model answering in a single pass is investigated, finding several mechanisms behind the gains.

Victor Gao, Vida Khosrowshahi, Ali Khosrowshahi et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.