It is shown that multi-byte prediction strikes a Pareto-optimal trade-off across multiple generative tasks, instruction following, question answering, summarization, and machine translation, achieving the best trade-off between performance and inference throughput.
Abraham Toluwase Owodunni, C. Okocha, Christan Grant et al.· 0 citations
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased usage, and instruction-guided models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage.
C. Okocha, Christan Grant, Zoey Liu· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.