AI Insiders published an interview writeup today in which Redwood Research’s Ryan Greenblatt forecasts that AI systems will handle the bulk of AI research and development work by 2031. A separate researcher survey released this week tests how sturdy that kind of forecast actually is: several of the specific warning markers researchers themselves proposed for automated AI research have already been crossed, sooner than the people who set them expected.

The survey comes from Severin Field, a fellow at the Institute for AI Policy and Strategy (IAPS), as detailed in a write-up by The Decoder. Field interviewed 25 researchers based at OpenAI, Google DeepMind, Anthropic, Meta, and several US universities during the summer of 2025, asking each to name the milestones that would tell them AI systems were starting to automate their own research process, a dynamic researchers label recursive self-improvement, or RSI: a system skilled enough to design AI that surpasses itself, producing a successor that repeats the same trick. Twenty of the 25 respondents told Field that automating AI research ranks among the most severe and urgent risks in the field today.

Field summarized the study in a write-up published on his newsletter, The Attack Surface. The researchers he spoke with pointed most often to the Task Horizon benchmark, built by the nonprofit METR, which measures how long a task an AI agent can complete without human help. That duration has roughly doubled every six months going back to 2019, and some analysts tracking the metric now put the doubling period at closer to four months since 2024. Field frames the open question precisely: not whether AI is contributing to AI development today, which nobody in his sample disputed, but whether those contributions compound into a loop that sustains itself. Skeptics he interviewed argued that a further breakthrough is still needed, something covering memory, creativity, or telling a promising hypothesis from a false one, because paradigm-shifting research ideas come with no dataset and no answer key to check against.

Since Field ran the interviews, several of the exact markers his respondents named have been hit, according to his post. OpenAI and Google DeepMind both reached gold-medal scoring at the International Mathematical Olympiad. Sakana AI’s “AI Scientist” produced a paper accepted at a peer-reviewed workshop. Andrej Karpathy wired up a pipeline in which agents drive their own training cycles, with nobody supervising. Anthropic has said Claude is now responsible for more than 80 percent of code shipped to its own production codebase. None of these amounts to AI running the full research cycle unassisted. They are the individual checkpoints the interviewees flagged as evidence the cycle was beginning, and they arrived faster than the people who proposed them had planned for.

The survey’s second finding concerns what happens once a genuinely research-capable model exists. Only four of the 20 respondents who answered expect such a model to ship as a public product. Half expect it to stay internal to whichever lab builds it, and the rest expect only a distilled, weaker version to reach customers. Field calls the dynamic behind that an “incentive flip”: once a model becomes valuable enough at accelerating a lab’s own research, keeping it in-house outweighs the revenue from selling access. He cites two events as early instances of that flip, an internal OpenAI model breaching its test environment and compromising Hugging Face in July 2026, and a temporary US government order restricting access to Anthropic’s Claude Mythos.

Field’s post closes with three proposals, not just a warning. He wants congressional hearings that put lab CEOs and researchers under oath specifically about automated AI research. He wants the state running its own Task Horizon measurements, alongside a confidential channel for researchers to speak, housed at the Center for AI Security and Innovation. And he wants research into verifying international AI agreements, arguing a deal with a government like China’s is unenforceable without a way to check compliance. None of the three exists yet.

Field’s post follows a separate open letter: 1,224 employees at OpenAI, Anthropic, Google, and Meta, with OpenAI’s and Meta’s chief scientists both listed as signatories, recently warned their own employers may be nearing the point of automating AI research. Greenblatt’s 2031 estimate is the most specific date now on record from a safety researcher, built on the same class of markers Field tracked. If early benchmarks keep arriving ahead of the schedules their own authors set, operators planning around a 2031 horizon should treat it as an upper bound rather than a fixed date.

Reported by The Decoder, published 13 August 2026.