Tom Davidson, a researcher at the AI-focused nonprofit Forethought, published an essay this week making a specific, testable claim: shortages of training data will not be what keeps AI from automating most of the economy.
Davidson’s target is a scenario some AI safety researchers call a “software intelligence explosion,” in which frontier labs use AI systems that have reached parity with human AI researchers to accelerate their own development from inside a data center. His essay works through several data bottlenecks that skeptics raise, and argues each one causes delay rather than a hard stop.
The first move is definitional. Davidson concedes that today’s AI models are extremely data hungry, needing far more examples than a person would to learn a task. But he argues the actual mechanism of a self-accelerating buildout is a shift toward sample-efficient learning algorithms, not brute-force scaling of datasets. He draws an analogy to DeepMind’s AlphaZero, which mastered chess and Go using the same underlying algorithm despite the two games sharing no training data.
That analogy is the load-bearing assumption in the whole piece, and it is the one a skeptic should press hardest. Chess and Go are fully observable, deterministic games with clean reward signals. Diagnosing a factory malfunction or negotiating a contract is not. Davidson’s essay itself concedes that neural networks trained on one domain generalize poorly to another. His entire forecast depends on learning algorithms behaving very differently from the models they produce, a distinction that is plausible but far from demonstrated at the scale he describes.
Where Davidson does find real friction, he still argues it produces delay rather than a wall. He describes a “paradigm tax”: once AI starts inventing techniques beyond current methods such as the Transformer architecture, there will be no equivalent trove of tutorials, forum threads, and example code for it to study, so each new advance has to be documented by AI itself before it can be mastered. He makes a parallel argument about data quality, citing reporting that OpenAI’s Project Mercury paid roughly 100 former bankers to construct sample financial models for training, at rates near $150 hourly, alongside a similar shift at Scale AI and Mercor toward credentialed experts. Once training data quality exceeds anything humans have produced, he argues, systems must build better examples from nothing instead of simply copying existing experts, and that process is slower.
By Davidson’s own estimate, these frictions could stretch a software intelligence explosion to roughly twice its otherwise-expected length, not prevent it. The larger uncertainty he flags for what comes after is whether companies will hand over proprietary data at all, since for many firms that data is the entire competitive moat. He predicts most will eventually sell access, priced above their own future profits, once one competitor in an industry breaks ranks.
None of this is empirical evidence. It is a forecasting model built on assumptions about how learning algorithms generalize and how fast AI can improve its own training data, and Forethought has not published results testing those assumptions against real systems. Teams setting AI roadmaps around a gradual, data-limited pace of progress should read this essay as a prompt to stress-test that assumption, not as a settled timeline.
Forethought published this essay by Tom Davidson on September 9, 2026.