Skip to content

Category

small language model

813 papers

#small language model Review Open access Aug 2026

Challenges and opportunities in type III secretion system effector prediction.

The conceptual evolution of T3SE prediction is reviewed, persistent limitations and sources of bias are highlighted, and open questions that must be addressed are outlined to enable robust, interpretable and ecologically inclusive prediction of T3SEs, pointing towards the need for centralized, user-friendly platforms that integrate diverse biological signals into transparent, ranked outputs suitable for experimental validation.

Iva Rosić, Ivan Nikolić · 0 citations
#small language model Preprint Aug 2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

ProViP is proposed, a training-free progressive visual token pruning framework that removes redundant visual tokens based on the embedding similarity of input tokens before reasoning of the LLM backbone, and then prunes tokens during reasoning via head-aware pruning.

Chaofang Ma, Lin Jiang, Carol Jingyi Li et al. · 0 citations
#small language model Preprint Aug 2026

Short Horizons and Sparse Concepts: a Mathematical View of the Readout in the J-lens

The Jacobian lens (J-lens) has been proposed as a way to read verbalizable representations from language models. However, its principle and meaning lack a detailed and theoretical discussion. We provide a mathematical view of this interpretation and of its assumed causal structure. Besides treating the J-lens as a heuristic probe, we further regard it as a first-order causal transfer operator from intermediate activations to expected future readouts. We study the Jacobian matrix as the optimal local linear approximation of the downstream mapping, analyze its global approximation behavior and bias, and identify its mathematical meaning as an expectation over anticipated future readouts. Further analysis of the Jacobian energy distribution reveals that its causal geometry is highly sparse. The energy decays with depth, concentrates in an extremely small proportion, and decomposes into diagonal pathways and specific critical positions. This decomposition further resolves the expectation of the J-lens over future outputs into short-horizon and sparse concept predictions, providing a more intuitive attribution and explanation for the ability of the J-lens to visualize concepts during the thinking process. Based on the theory, we propose a simple but effective improvement strategy and decoupling method for the J-lens, which significantly enhances the ability of the J-lens to read out correct intermediate concepts.

Shi-Qi Yan, Kai-Xuan Ding, Chao-Hong Tan et al. · 0 citations
#protein folding Review Aug 2026

Emerging experimental and computational methods for studying redox-regulated structural transitions.

A comparative overview of dual experimental and computational advancements is provided and how the integration of generative diffusion models could facilitate the real-time simulation of conditional, multi-state structural ensembles across the redox proteome is highlighted.

T. Rass, Dana Reichmann, Gábor Erdős · 0 citations
#small language model Open access Aug 2026

CARE: conflict-aware regulation and evolution for continual test-time open-vocabulary semantic segmentation

A continual test-time adaptation framework that updates only a small subset of model parameters that consistently outperforms existing TTA and OVSS adaptation baselines across multiple datasets and corruptions, while maintaining stable performance over extended and cross-corruption continual streams without introducing additional trainable modules.

Fen Luo, Sen Li, Zhanqiang Huo · 0 citations
#small language model Open access Aug 2026

What students ask matters: LLM interaction depth, task quality, and immediate recall in higher education

The findings indicate a dissociation between performance quality and short-term recall in LLM-supported study, which aligns with cognitive-psychology evidence that elaboration improves comprehension while retrieval practice consolidates retention.

V. Tsiligkiris · 0 citations
#small language model Preprint Aug 2026

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.

Siyao Yan, Bo Han, Jisheng Dang et al. · 0 citations
#small language model Preprint Aug 2026

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems, and the 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form.

Apodex Team B. An, B. Li, B. Wang et al. · 1 citation
#small language model Preprint Aug 2026

MGQL: An Executable, Small-Step Semantics of GQL

MGQL is presented, the first mechanized, small-step operational semantics for a substantial read-only fragment of GQL that is grounded in the ISO/IEC 39075 standard, and it is proved that the type system is sound, ensuring an end-to-end guarantee of well-formed queries yielding results that conform to their declared schemas.

Aditya Thimmaiah, Tongtong Lin, Milos Gligoric · 0 citations
#small language model Preprint Aug 2026

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

The role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation is investigated, and a key insight is revealed: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach.

Qi Ye, Zhi-Yuan Gu, Jingjie Xia et al. · 0 citations
#small language model Preprint Aug 2026

Dataset Scarcity Limits Robust Evaluation of Multilingual Embedding Models: A Case Study of Slavic Languages

A two-dimensional framework, specifically tailored for analyzing multilingual embedding benchmarks under dataset scarcity, is proposed and applied on the Slavic-language subset of the MTEB benchmark, revealing severe benchmark sparsity.

Ana Gjorgjevikj, B. Seljak, T. Eftimov · 0 citations
#small language model Preprint Aug 2026

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

Ternary multiplicative adaptation is proposed, which represents discrete updates of ternary weights such as sign flips or zeroing through a low-rank Kronecker factorization into two small ternary matrices applied element-wise to ternary weights.

Alexandru-Dragos Manolache, Yun-qiang Li, Jan van Gemert · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.