A language model can get a simple sum right, forget how to do it a few steps later, and then get it right again, according to a paper posted on alphaXiv by Jiaxin Wen, Zhengxuan Wu, Dawn Song, and Lijie Chen. The authors call the behavior mode-hopping: the model switches between copying a tempting pattern in the prompt and working out what the task actually asks.
The clearest example uses OLMo3-32B, an open model from the Allen Institute for AI. In the researchers’ “answer+1” test, the examples show the sequence 1, 2, 3, so the lazy continuation is 4. The correct arithmetic answer is 8. At 2.17 trillion training tokens the model chose the correct answer 81% of the time. At 2.19 trillion it managed 0%. At 2.21 trillion it was back at 81.7%.
That matters because the usual dashboards would have missed it. Teams watch training loss, a running measure of prediction error, and scores on familiar benchmarks. According to the authors, those curves stay smooth and stable while behavior underneath them swings. The authors stress that the arithmetic reversal is a single task-specific example, so it says nothing about whether other abilities flip in the same pattern.
The team built a set of six quick tests that need only a handful of examples, plus two that involve fine-tuning. They probe whether a model follows surface habits or infers the task. One pits flipped labels against familiar sentiment patterns. Another separates truth from merely sounding true. A third checks intuitive wrong answers against careful reasoning, and a fourth asks whether scattered facts hang together as a coherent persona.
The authors also tried to rule out boring explanations. A well-known 2023 critique argued that apparent sudden jumps in ability can be an artifact of harsh scoring. Wen and colleagues report that mode-hopping shows up under both strict accuracy and a softer probability measure, and whether they plot against tokens or raw compute. A single training step on random data barely moved the results, even at an aggressive learning rate. Averaging five checkpoints smoothed the swings but did not remove them.
The practical finding is about picking where to stop. The researchers compared two OLMo3-32B snapshots, taken at 4.5 trillion and 4.9 trillion tokens, and then fine-tuned both. After math fine-tuning, the earlier snapshot scored 36.3% on the GPQA science benchmark against 29.8% for the later one. After general post-training, it resisted prefilling attacks (where an attacker starts the model’s reply for it) 53% of the time versus 21%. The paper cautions that this does not show earlier checkpoints win in general, only that an intermediate one can be the better starting point under these setups.
A small experiment pointed at a possible remedy. Continuing training from an intermediate checkpoint on data chosen to favor pattern-matching, or to favor genuine generalization, pushed the model steadily toward each behavior. The run on random data kept hopping. The authors flag this as preliminary, covering one test at small scale.
The explanation they offer is a hypothesis: shallow shortcuts and more general circuits compete for limited capacity inside the model, and the winner changes from checkpoint to checkpoint. The experiments measure behavior, not the circuits themselves, so the paper stops short of naming what actually drives the flips. Results also differ from one dataset to the next, although correlations run higher in larger models.
Labs have long treated the final checkpoint as the best one by default. If these findings hold beyond OLMo3 and the Apertus models the authors also tracked, a cheap behavioral check run during pre-training could decide which snapshot goes into expensive post-training, and the authors have released their test suite on GitHub so others can try.
Reported from the paper “Generalization Dynamics of LM Pre-training” by Jiaxin Wen, Zhengxuan Wu, Dawn Song, and Lijie Chen, read on alphaXiv. The source carries no publication date.