John Schulman, the chief scientist at Thinking Machines and a co-founder of OpenAI, told podcast host Dwarkesh Patel that AI systems could dominate top human experts across every field of cognitive work within three to four years. Charlie O’Neill, head of model training at Baseten, put the same milestone at five to ten years. Beren Millidge, chief technology officer at the open-source model lab Zyphra, landed in between: roughly five years for the domains labs already prioritize, longer for everything else.
That is a threefold spread on the single question that determines how fast AI research compounds on itself, from three people who train or build frontier-scale systems for a living. The disagreement is not really about a date. It is about whether the current recipe, transformers trained with reinforcement learning on ever-larger environments, can discover the next breakthrough on its own, or whether progress needs a discontinuity nobody has found yet.
O’Neill argued for the second view. He compared the situation to Moore’s Law, which held as a straight line for decades only because engineers kept finding discrete new tricks to extend it each time the old one hit diminishing returns. Pretraining scaling hit that wall and reinforcement learning replaced it. “I don’t think, if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you’re running, is necessarily capable of discovering” the next such trick, he said, if that trick lies far enough outside what gradient descent and neural networks can reach.
Millidge pushed back on the premise that a persistent gap between simulated and real-world performance would stall progress indefinitely. “We do actually see this kind of generalization even from RL in practice already,” he said, arguing that if a system fails to generalize widely from reinforcement learning it is more likely a training problem than a hard ceiling.
Schulman’s version of skepticism was different again: not a missing discontinuity but a recurring bottleneck. He described a pattern where “a new model comes out and people are blown away and they’re like, ‘This is it. This is AGI.’ But then they use it a bit, and it starts to feel dumb after a month or so.” Code generation is already fast enough to make researchers many times more productive on narrow tasks, he said, yet that speed has not translated into 100x overall research output, because judgment, taste, and self-checking remain the limiting factors.
The three numbers reflect three different bets on which bottleneck breaks first. Schulman’s three-to-four-year estimate assumes the hard parts (judgment, longer-horizon learning) get solved on roughly the same curve as everything else. O’Neill’s five-to-ten-year estimate assumes automating AI research specifically is close to what he called “ASI-complete,” gated by problems like needing more context than a million tokens can hold. Millidge’s answer split the difference by domain: parity soon wherever compute and attention are already concentrated, and a long tail of expertise nobody has bothered to build training environments for.
None of the three treated the outcome as inevitable on any of these timelines. Anyone building a product roadmap on “AI matches expert humans by year X” is choosing among three specific, disagreeing bets made by people inside labs actively running that experiment, not a converging consensus. The domains where these researchers expect the earliest gains, coding and other data-rich, text-native fields, are the ones worth stress-testing against real workflows now, since even the most bullish of the three sees that as a three-year floor before broader claims are testable.
Reported from Dwarkesh Patel’s podcast interview with John Schulman, Beren Millidge and Charlie O’Neill, published September 11, 2026.