Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on it...
Shu-Xiao Xie, Shu-Yang Xie, Yuan Cao et al.· 0 citations
Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is...
Shu-Xiao Xie, Shu-Yang Xie, De-Zhi Ran et al.· 0 citations
Algorithm realization becomes an independent low-precision axis with a global $\Phi$ optimum and measured fp8 relevance: an independent low-precision axis with a global $\Phi$ optimum and measured fp8 relevance.
Shu-Xiao Xie, Shu-Yang Xie, Yuan Cao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.