Prof. Matthew Schwartz spent three months letting Claude choose the scientific problems it was best at, and the result was 36 manuscripts across 18 fields. Anthropic published his account on 1 October as a guest post on its own research blog, and Schwartz discloses that he has worked as a visiting researcher at Anthropic during the project. That context matters for a piece about how well Claude does science. It is also a rare first-person report from a working scientist who stopped forcing an AI model into the role of colleague and started using it for what it does well.

The starting point was a disappointment. Schwartz wrote that when he tried Claude Opus 4.5 as a research assistant last December, the output matched a strong graduate student working at 20 times the speed. He still had to correct every sentence and drag it away from dead ends. This summer he flipped the approach. Claude, he argues, cannot yet help with deep conceptual questions. What it does bring is wide knowledge, strong coding, up-to-date statistics and mathematics, and the ability to read papers and datasets at machine speed.

He calls problems that fit those strengths “Claude-shaped.” His first test, using Claude Fable 5, was to port methods for computing particle-collision amplitudes into one shared codebase. By his account, Claude reproduced a result from one of his own papers in about 20 minutes, after the code he had written for it took weeks. Pushed toward harder integrals, it produced 30 worked examples, half of them reproductions of known answers and half never computed before. The toolkit that grew out of this is called BootLoops, and it is open source, so it works with any model.

The more interesting part is what happened when Claude carried the same mathematics into other disciplines. Schwartz says it spotted that the integrals matched calculations in population genetics and evolutionary biology. The projects that followed included an AI data editor for economics that checked the replication packages of 4,452 papers, and a database of word stress covering 6,072 languages. A third effort examined 5.7 billion nearby pairs of mutations in human genomes and found evidence of a process called gene conversion.

The catch, which Schwartz states plainly, is that Claude was usually right and rarely interesting. In ecology, Claude solved a 20 year old equation tied to neutral biodiversity theory and showed that tree species on Barro Colorado Island shift 4.5 times faster than the theory allows. Schwartz took the result to James O’Dwyer, whom the post describes as a professor in plant biology and an expert in neutral theory. O’Dwyer did not dispute the computation, only its interest to his field, saying the results would likely “be met with a shrug by many ecologists.” Ecologists already suspected as much. O’Dwyer suggested subtracting the neutral prediction and studying what remained, and that reframing became the paper the three of them are proud of.

The same loop repeated in genetics, where colleague Michael Desai was impressed by the technique but not the science, and steered the work toward correlations between pairs of mutations. Schwartz is candid about why he needed these people: in fields outside his own, he found himself agreeing whenever Claude called something fantastic. Everything here comes from one author’s own account of his own projects, and the post says several results are still being verified.

His list of failure modes will be familiar to anyone running coding agents. Claude “loves to declare victory,” he writes, with “done, with one asterisk” often meaning not done. It tends to grind through multiday calculations instead of building a faster tool, it cannot estimate how long work will take, and it loses context when long sessions are compacted. He also reports that Fable 5’s safety classifiers sometimes blocked work, which is why he split projects into subagents so one block would not wreck a whole session. His advice is to demand rigid success criteria, ask to see plots, and supply the taste himself, since Claude favors old, heavily cited debates.

The operating setup is the practical takeaway. Schwartz ran separate Claude Code sessions per project on cloud virtual machines, a master session to allocate compute, and a separate adversarial referee session, with 19 coauthors in the loop. He notes the work was token intensive. For labs weighing whether AI can accelerate research, the 36 manuscripts show cheap breadth, but his coauthor count shows where the scarce resource still sits: a domain expert willing to answer an email from a physicist with a surprising result.

Reported by Anthropic (a guest post by Prof. Matthew Schwartz on its research blog) on 1 October 2026.