Tom Zahavy, a research scientist at Google DeepMind, argues in a new position paper that large language models are missing a specific cognitive move required for genuine scientific breakthroughs, not just faster ones. The Decoder reported on the paper, titled “LLMs can’t jump,” on July 30, 2026. Zahavy’s claim matters for anyone betting on AI to compress decades of research into months.

Zahavy borrows philosopher Charles Sanders Peirce’s three-way split of reasoning: deduction, working forward from fixed rules, induction, spotting patterns across many cases, and abduction, inventing an explanation for something unexpected. He grants that models already handle the first two well, pointing to systems like AlphaProof, Gemini, and GPT-5 clearing gold-medal scores on International Mathematical Olympiad problems. The gap, he argues, sits inside abduction itself. Ordinary abduction means selecting the strongest option from a menu of explanations the field already recognizes, something Zahavy says models can do. What they cannot do, in his framing, is manipulative abduction: coining an explanation that has no prior term or precedent to select from in the first place.

Einstein is his test case. When Einstein worked out general relativity, Newtonian physics fit the data almost perfectly, and astronomers had already accounted for the one visible anomaly, a small drift in Mercury’s orbit, by proposing a hidden planet named Vulcan. A system trained to minimize prediction error against reality would have had no error to chase, Zahavy argues, and so would have kept patching the existing model rather than replacing it. Confirming evidence for relativity did not arrive until years after Einstein published the theory.

What produced the leap instead, in Zahavy’s account, was a grounded thought experiment: Einstein pictured a physicist riding inside a falling elevator who feels no gravity, and Archimedes reportedly reached his buoyancy principle through the physical sensation of water rising around him in a bath. Language models, trained purely on text, have no equivalent sensory channel, a point he ties to philosopher John Searle’s Chinese Room, where a person manipulates symbols without understanding them. He extends the point to existing automated-science tools: Sakana’s AI Scientist recombines concepts that already exist, and DeepMind’s own AlphaEvolve needs a shrinking error signal to optimize against, something Einstein never had.

His proposed way out is a world model, defined here as a system trained to carry a predictive internal representation of how physical reality behaves rather than a representation of text describing that reality. Zahavy separates a video generator like Veo, which predicts the statistically likely next frame, from an action-controllable system like Genie, which lets an agent intervene in a simulation and test counterfactuals, such as cutting a cable to see what falls. That kind of synthetic lab, he argues, could supply the feedback loop manipulative abduction needs.

The argument has a real counterweight. In a separate article AI Insiders is publishing today, Anthropic said one of its models produced genuinely new cryptanalytic attacks, output that looks like more than recombining what already existed in the training data. That tension is worth sitting with rather than waving away. Zahavy’s framework has an answer available: cryptanalysis is a closed, rule-governed search space in a way physics was not for Einstein, so a model can find a novel combination inside defined rules without needing an embodied intuition about the world. Whether that distinction holds for every case claimed as a language-model breakthrough is the actual open question, and it is not settled by either paper alone.

Zahavy’s case is strongest as a description of where today’s automated-science tools plateau, weaker as a permanent ceiling on the category. A training signal problem is not the same claim as an architectural impossibility, and the paper argues the former while sometimes reading like the latter.

Teams evaluating AI for research and development should stop asking whether a model can generate hypotheses and start asking what kind of error signal, if any, it was trained against, since that distinction is what separates incremental optimization from a genuine shot at discovery.

The Decoder’s Maximilian Schreiner reported on Tom Zahavy’s Google DeepMind position paper “LLMs can’t jump” on July 30, 2026.