Large language models (LLMs) require increasing memory capacity to accommodate growing model weights and KV caches. High-Bandwidth Flash (HBF) offers high memory density and aggregate read bandwidth through massive plane-level parallelism, making it an attractive option for LLM serving. However, serving LLMs entirely f...
Shu-Zhang Zhong, Wei-Kai Xu, Yi-Fan Zhou et al.· 1 citation
NAMOH, an architecture-native sparse attention mechanism that activates only its assigned tokens and performs causal attention within this subsequence, is introduced, and it is hoped this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.
HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...
Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al.· IEEE Transactions on Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.