Skip to content

Author

Xiang-Yu Zhang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Looped Transformers as Optimizers

Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transit...

Yu-Long Huang, Chen Jiang, Zhan-Peng Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

KV-Invariant Transformer Expansion (KITE) is introduced, a scaling paradigm that trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV.

Zhi-Heng Hu, Yi-Xun Wei, Jian Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.