Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models have likewise highlighted the value of scaling test-time computation through longer computation trajectories. However, the principles for designing effective loop transit...
Yu-Long Huang, Chen Jiang, Zhan-Peng Zhou et al.· 0 citations
LeapQuant is proposed, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization and substantially reduces memory and compute costs during inference.
HyperTransfer is proposed, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself, to extend the framework to non-scale-invariant networks.
Jinghui Yuan, Hong-Tao Zhang, Jade Zou et al.· 0 citations
This work proposes Software-Interposed Datapath (SID), an efficient, software-flexible, and low-cost solution for rack-scale interconnects, particularly optimized for fine-grained memory access.
Chenxingyu Zhao, Yibo Wu, Hong-Tao Zhang et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.