Skip to content

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

KV-Invariant Transformer Expansion (KITE) is introduced, a scaling paradigm that trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV.

Abstract

Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves this goal. It trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV. Consequently, during inference, prefilling KV only relies on the smaller part of the model, so the inference costs are saved. As a concrete instantiation, we present Step Scale Transformer (SST), a two-tower decoder in which one tower produces KV and the other reads them. At comparable cumulative training compute, SST, a 67B MoE model with 2.15B active body parameters per decode token, achieves lower training loss than 47B and 63B MoE Transformers with 1.48B and 2.02B active body parameters, respectively, while reducing estimated inference cost by 6.7% and 31.6%.

View source

Similar papers

#machine learning Preprint Oct 2026

Learning Rate Transfer for Hybrid Transformer-SSM Architectures

We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at inf...

Jimin Seo, Gyubok Lee, Yeonsik Jo et al. · 0 citations
#natural language process... Preprint Aug 2026

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

Reduced Matrix Multiplication is proposed, a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights, and it is shown that the same principle extends to multimodal vision-language inferenc...

Zi-Xuan Lan, Yan-Hong Li, Jia-Wei Zhou · 0 citations
#artificial intelligence Preprint Sep 2026

LoopICL: Looping a single transformer block to solve tabular tasks

Tabular foundation models using in-context learning have recently surpassed gradient-boosted trees on predictive tabular tasks. However, recent mechanistic insights suggest that parameters in these models are largely redundant. We introduce LoopICL, a looped transformer whose core design decouples parameter count from...

Amir Rezaei Balef, Katharina Eggensperger · 0 citations
Preprint Aug 2026

Rethinking Expressivity and Efficiency in Test-Time Training

Under the standard approximation of taking gradients at the chunk-start weights, a closed-form state transition is derived that exactly reproduces the chunk-end fast-weight and momentum states of the per-token recurrence of Test-Time Training.

Zeyun Zhong, Joya Chen, Manuel Martín et al. · 1 citation
Preprint Aug 2026

Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation

Circuit Fine-Tuning is introduced, a compute-efficient framework that uses circuit discovery---conventionally used to explain trained models---to select modules for fine-tuning before training to isolate the response of the backbone to the target distribution rather than the preferences of a particular classifier.

Uri Z. Kialy, Gil Ben-Artzi · 0 citations
Preprint Aug 2026

MixFormer: Linear Transformer with Mixture of Memory Experts

State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in long-context modeling. However, existing SSMs suffer from limited input adaptivity and constrained memory capacity, leading to information loss when modeling ultra-long se...

Yu Guo, Lei Duan · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.