A small note on a machine-learning idea

Grokking

This is Brandan McMurtrie’s own page. It is not an official Grok, xAI, or Elon Musk site. The domain name is a nod, in good faith, to that project — nothing here speaks for it.

Sometimes a small neural network looks finished too early. It gets the training examples right and fails on anything new, which is what memorizing looks like. Then, long after that, and with no new data, the answers on unseen examples jump from roughly chance to nearly perfect. Researchers call that late jump grokking.

Where it showed up

The name, in this sense, comes from a 2022 paper by Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra: “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets” (arXiv:2201.02177).

They did not start with photographs or language. They used tiny, made-up tables of equations, like “a ∘ b = c”, where each symbol is just a token with no built-in meaning. The network has to fill in the missing rows of a puzzle it has only partly seen. One striking case is division modulo 97, with half the equations held out. Training accuracy was already close to perfect in under a thousand optimization steps. Validation accuracy stayed near chance until around a hundred thousand steps, and only approached perfect near a million.

People cared because the gap is so clean. On ordinary datasets, “it memorized” and “it understood” blur together. Here you can watch a network sit in the memorized state for a very long time and then generalize anyway. The same paper found that smaller training sets need much more optimization before that happens, and that weight decay helped more than most of the other tweaks they tried.

A closer look, on one toy

Later work opened one of these toys and looked at the weights. In “Progress measures for grokking via mechanistic interpretability” (arXiv:2301.05217; Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt), small transformers trained on modular addition learned to treat numbers as points on a circle and add them with sines and cosines. On that task the sudden jump in test accuracy came after a generalizing circuit was already forming. What looked like a single moment was, in their telling, three stretches: memorize the training pairs, slowly build the real algorithm, then let weight decay clear out the memorized bits so the algorithm can show through.

What is still open

Both papers are about small models and made-up tasks. It is not settled how often the same story appears in large models trained on real data, or how to tell in advance when the late jump will come. Weight decay mattered a lot in these experiments; that does not mean it is the whole explanation everywhere. Grokking is a sharp example of generalization arriving late. It is not yet a finished theory of why networks generalize.