Anthropic has started publishing internal metrics on how much of its own AI research is now run by Claude rather than by people, and by its own accounting the model has crossed a notable threshold. As of August 2026, Claude “leads” 26% of the company’s AI research and development work, meaning it can carry a task from a high-level instruction to a finished result while a human supervises rather than steers. More than 90% of that R&D work now involves AI at or above a “collaborates” level, though Anthropic says no measured task has reached full autonomy.

These figures come from a new tool the company calls the Anthropic R&D Automation Index, built by mapping roughly 15,000 internal tasks into 542 categories and having Claude itself help rate how automated each one is. That is a self-graded scale: Anthropic checked it against human raters and reports 59% exact agreement, with model and human scores landing within one level of each other 97% of the time. The company is candid that this is not independent verification, and says it plans to embed outside evaluators with access comparable to its internal risk teams.

The automation numbers arrive alongside two other releases: how closely Anthropic watches the agents doing that work, and how much computing power goes to safety research versus everything else. On oversight, Anthropic says roughly 30,000 agents were active on its main internal platform at any given moment in August, with every action passing through a real-time monitor before execution. Out of more than a billion agent decisions logged that month, 0.002%, about one in 47,000, were blocked. A separate after-the-fact review process flags around 100,000 transcripts a week, of which roughly 50 get escalated to a human.

On compute, Anthropic examined one week in July 2026 and found about 6% of the processing power spent on AI research and development went to safety work, rising to 12% when the work involved AI itself doing the research. The company frames this as a deliberately conservative estimate that excludes classifiers and other safety infrastructure that runs alongside production systems.

Anthropic is positioning all three measurements as evidence for a broader argument: that the industry needs shared, third-party-checked ways to track the pace of AI development, an idea chief executive Dario Amodei has pushed publicly as a case for coordinated pacing among labs. The company says any of its measures, the automation index, the monitoring statistics, or the compute split, could be replicated by competitors using a public methodology.

That pitch has an obvious tension built in. Anthropic is both the subject of these measurements and the only party currently producing them, using its own models to grade its own automation levels and classify its own compute. The company acknowledges as much, noting that a “judge” model checking a company’s systems could share the same blind spots as the system it is judging, which is precisely why it says outside verification needs to start soon rather than later.

For any lab or regulator weighing how to track AI automation, Anthropic’s numbers set an early public benchmark, but one that only holds up once a rival lab or a truly independent evaluator publishes a comparable figure using the same yardstick.

Anthropic, “Measurements for understanding the pace of AI development inside frontier labs,” published on the Anthropic Institute site.