KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
KV-Invariant Transformer Expansion (KITE) is introduced, a scaling paradigm that trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV.
Zhi-Heng Hu, Yi-Xun Wei, Jian Zhou et al.
· 0 citations