A method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module, allowing pretrained models to operate on compressed inputs without architectural changes or re-training.
Zi-Xuan Lan, Jessica H. Yang, Yan-Hong Li et al.· 0 citations
Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoke...
Reduced Matrix Multiplication is proposed, a training-free, input-adaptive inference method that reduces Transformer matrix products by selecting informative slices along their contraction dimensions, without modifying model weights, and it is shown that the same principle extends to multimodal vision-language inferenc...