Skip to content

Author

Daxin Jiang

We have 3 of 47 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

KV-Invariant Transformer Expansion (KITE) is introduced, a scaling paradigm that trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV.

Zhi-Heng Hu, Yi-Xun Wei, Jian Zhou et al. · 0 citations
2025

Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs

This work introduces Farseer, a novel and refined scaling law offering enhanced predictive accuracy across scales, and provides new insights into optimal compute allocation, better reflecting the nuanced demands of modern LLM training.

Houyi Li, Wen-Zheng Zheng, Qiufeng Wang et al. · 4 citations · ⚡1
Preprint Aug 2026

Scheduling Mixed RL Rollouts Beyond Prefix Locality

MISA-T, a routing-layer admission policy for mixed rollout serving that combines adaptive session admission, workload-aware KV-capacity allocation, and residency-time-aware KV accounting, is presented.

Zetao Hong, Song Yuan, Yuanhao Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.