Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head
NAMOH, an architecture-native sparse attention mechanism that activates only its assigned tokens and performs causal attention within this subsequence, is introduced, and it is hoped this work offers a new path for scaling attention, with parameter scaling directly enabling context scaling.
Zi-Zhuo Fu, Run-Sheng Wang, Meng Li
· 0 citations