A one-day, fully pre-registered study of grokking interventions on modular addition and small language models. Main results: (1) published grokking 'accelerators' largely reduce to artifacts of untuned baselines and first-crossing metrics; (2) an auxiliary per-example memory races the network for training CE and, via dead gradients plus weight decay, erases even an already-found generalising rule (cycle-vs-death dose structure; two independent erasers; on LM continued training the dominant eraser of unreinforced knowledge is interference, not weight decay); (3) at fixed data size, training-set structure decides grokking (with a pre-registered character-sum predictor) and transfers to language as a conditional two-axis law (coverage and window freshness) with an ordinal data predictor; (4) the ReLU/GELU gap is a training-budget effect while the final LayerNorm matters for both activations. Every claim was staked in a ledger before computing (176 stakes, 158 outcomes); the ledger is attached. Companion code and data repositories: https://github.com/iwasborninbali/grokking-honest-ruler, https://github.com/iwasborninbali/grokking-race-eraser, https://github.com/iwasborninbali/grokking-data-law, https://github.com/iwasborninbali/grokking-activation-ln. The experiments, analysis and text were produced by autonomous AI agents under the direction of the author of record, who takes responsibility for the content.
Aleksei Kudriashov, Claude Fable 5 (Anthropic) — autonomous research agent, ChatGPT gpt5.6-sol (OpenAI) — deep research and review· Zenodo (CERN European Organi...· 0 citations
A one-day, fully pre-registered study of grokking interventions on modular addition and small language models. Main results: (1) published grokking 'accelerators' largely reduce to artifacts of untuned baselines and first-crossing metrics; (2) an auxiliary per-example memory races the network for training CE and, via dead gradients plus weight decay, erases even an already-found generalising rule (cycle-vs-death dose structure; two independent erasers; on LM continued training the dominant eraser of unreinforced knowledge is interference, not weight decay); (3) at fixed data size, training-set structure decides grokking (with a pre-registered character-sum predictor) and transfers to language as a conditional two-axis law (coverage and window freshness) with an ordinal data predictor; (4) the ReLU/GELU gap is a training-budget effect while the final LayerNorm matters for both activations. Every claim was staked in a ledger before computing (176 stakes, 158 outcomes); the ledger is attached. Companion code and data repositories: https://github.com/iwasborninbali/grokking-honest-ruler, https://github.com/iwasborninbali/grokking-race-eraser, https://github.com/iwasborninbali/grokking-data-law, https://github.com/iwasborninbali/grokking-activation-ln. The experiments, analysis and text were produced by autonomous AI agents under the direction of the author of record, who takes responsibility for the content.
Aleksei Kudriashov, Claude Fable 5 (Anthropic) — autonomous research agent, ChatGPT gpt5.6-sol (OpenAI) — deep research and review· Zenodo (CERN European Organi...· 0 citations