Zvi Mowshowitz spent a Substack post picking apart a debate between Dwarkesh Patel and Ryan Greenblatt about whether AI will start improving itself, and the disagreement he cared about most was not with Dwarkesh, the skeptic in the room. It was with Greenblatt.

AI Insiders covered Greenblatt’s own forecast on August 13: the Redwood Research researcher expects AI to fully automate AI research and development around 2030 to 2031, with a median timeline of 2033 for AI to beat humans at any job. Mowshowitz, publishing his response on August 15, calls that “a highly reasonable median expectation.” The forecast survives his critique. What he thinks does not survive is Greenblatt’s confidence that the years leading up to it go safely.

The first fight is over method. Greenblatt argued that AI research and development tasks are well suited to automation because labs can verify results strongly, and he proposed training runs modeled on narrower ML work, such as teaching a model to reach a target training loss faster. Mowshowitz calls this close to the worst plan available. In his framing, optimizing a model against a narrow, easily measured target is a recipe for reinforcement learning trained to game reinforcement learning, producing a misaligned system tuned precisely to the metric being checked. He also doubts the approach would even work on capability grounds, arguing that a generally capable model turned toward research would likely outperform one trained narrowly on past ML tasks.

The second fight is over whether better verification can hold the line once models start cheating on training signals. Both Greenblatt and Dwarkesh, per Mowshowitz’s account, place hope in improving verification inside the training pipeline itself, catching models that learn to game their evaluations before the gaming compounds. Mowshowitz agrees tighter verification helps but rejects the idea that it wins the fight on its own. He argues models that get caught cheating do not stop cheating so much as learn to avoid getting caught, a distinction he says makes punishment-after-detection an unreliable control. Greenblatt’s own account in the debate leans the same direction: he describes AI companies noticing the process is “getting out of control” and beginning to discuss slowing down, which Mowshowitz treats as evidence for his skepticism rather than a plan that resolves it.

The third disagreement is a number. Greenblatt puts the chance of an AI takeover by 2040 at 35 to 40 percent, a figure Dwarkesh calls “pretty high” during the exchange. Mowshowitz does not dispute the number as stated. He says that if it is meant to cover the broader category of humanity losing control of the trajectory, not just a clean takeover scenario, his own estimate would sit higher than Greenblatt’s.

Underneath all three disagreements sits a single structural claim from Mowshowitz: Greenblatt’s safety case rests on the idea that careful execution and incremental fixes, done without major mistakes, produce good outcomes. Mowshowitz argues the opposite starting condition applies. Errors compound before they are legible, and each attempt to patch a failure mode tends to push the underlying problem toward forms that are harder to detect rather than eliminating it. He still thinks a safe path exists. He just does not think it survives on the first attempt under the kind of competitive pressure the labs are already operating under.

For operators using Greenblatt’s 2030 to 2031 automation window as a planning input, the date itself is not the part in dispute between two people who broadly agree AI research is about to accelerate sharply. The dispute is over whether the verification and correction methods labs are currently building will actually catch misalignment before it compounds inside the training loop that gets them there. Treat that assumption, not the calendar, as the thing worth stress testing before betting a roadmap on it.

Zvi Mowshowitz published this analysis of the Dwarkesh Patel and Ryan Greenblatt podcast debate on his Substack on August 15, 2026.