Preprint
Aug 2026
Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
The empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although the trend of stable gains is confirmed with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages.
Sofiia Riazhskykh, Nam Luu, Ondrej Bojar
· 0 citations