Anthropic published an account of an instance of its Claude model, internally called Fable 5.1, extending a particle physics calculation one step beyond the previous human record. The post was written by Matt von Hippel, a science writer and former theoretical physicist, and ran on Anthropic’s own research blog. Anthropic disclosed that it invited von Hippel to write the piece, paid him for his time, and had staff give feedback on drafts, while the opinions in the post remain his own. Two Anthropic physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, ran the calculation inside Claude Science, the company’s paid research platform, largely by giving Claude one prompt and then telling it to keep working for hours or days at a stretch.

The task grew out of a challenge von Hippel posted on his own blog in August, daring AI labs to extend a scattering-amplitude calculation, a formula physicists use to predict how subatomic particles interact. He picked N=4 super Yang-Mills, a simplified theory researchers use to stress-test calculation methods rather than describe real particles. Lance Dixon, a physicist at the SLAC National Accelerator Laboratory who has worked on this problem since the field began, had reached eight loops of complexity some years earlier by routing through a related quantity called a form factor. Claude’s result reached a ninth loop, and reached it directly, the step Dixon had expected would still need an indirect workaround.

Claude produced the nine-loop answer two separate ways: the direct bootstrap method, and the indirect form-factor route Dixon had used. The bootstrap version, coded in Python with the SymPy library, consumed about $100 of compute, equivalent to 96 CPUs running for a week. Anthropic said the actual total cost of the exercise came to a few hundred dollars. Only a paying customer running both methods end to end, rather than the internal team, would have faced a bill of one to two thousand dollars, mostly from how long Claude kept working.

The nine-loop result withstood independent scrutiny from Dixon, who confirmed it on his own. Anthropic gave him Claude usage credits for the work. In an addendum to the post, he wrote that the underlying setup was fragile enough that a single mistake anywhere in the calculation would have collapsed the whole result, and that Claude built its supporting code from scratch instead of drawing on his team’s unpublished tools.

The result turned out to be less exclusive than the original challenge implied. Within days, a separate group led by Song He at the Chinese Academy of Sciences reached most of the same nine-loop answer through its own methods, using GPT-6 for partial help rather than the largely unsupervised run Anthropic describes. Von Hippel, noting again that Anthropic had invited and paid him to write the account, concluded that Claude applied established techniques with more compute than earlier attempts had used, not a genuinely new method, and that the real finding was how much progress on this problem had simply gone unclaimed.

What should matter to anyone evaluating frontier models for research work is the operating pattern, not the physics trophy: days of largely autonomous effort on a narrow, checkable problem, reviewed by a named outside physicist rather than certified by Anthropic itself. Readers weighing that pattern should also weigh where the account came from: a blog post Anthropic commissioned and paid for, not an independent publication or a paper from Anthropic’s in-house physicists. Teams considering multi-day autonomous runs on their own technical backlogs now have a public example to benchmark against, funding relationship included.

Based on a guest post Anthropic commissioned from Matt von Hippel, a science writer and former theoretical physicist, published on the company’s research blog with an addendum from Lance Dixon of SLAC National Accelerator Laboratory, following his September 1 validation of the nine-loop result.