OpenAI framed the launch of GPT-6 Astra as the dawn of AGI, pointing to a 99.9 percent result on ARC-AGI-3 as the proof. We reported on September 5 that the same model, run through ARC Prize’s own standard harness rather than OpenAI’s provider adapter, scored 62.7 percent, a 37-point gap produced entirely by the scaffolding around the model rather than the model’s weights. That measurement is now established. What deserves scrutiny is the claim OpenAI built on top of it, and what has happened to the numbers since.

ARC Prize, the organization that built and administers the benchmark, has not endorsed the AGI conclusion. The foundation itself said flatly it is “not claiming that it is AGI,” while co-founder Mike Knoop added that “we lack evidence to call this AGI yet.” From here on, ARC Prize plans to show both harness scores together whenever it reports results, closing off the option of citing just one adapter-inflated figure in isolation. That is a direct rebuttal from the benchmark’s own authors to the interpretation OpenAI attached to their data.

AGI has no agreed operational definition across the industry, so any claim to have reached it is really a claim about one chosen benchmark under one chosen scaffold. That is why the harness detail is not a technicality. It is the entire basis for whether the claim can be evaluated at all, and ARC Prize’s decision to publish both figures going forward is an acknowledgment that a single headline number invites exactly this kind of overreach.

The comparison that spread publicly compounded the problem. GPT-6 Astra’s 99.9 percent (from the provider adapter) was set against GPT-5.6 Sol’s 7.8 percent (from the standard harness), a mismatch The Next Web says ran in its own earlier coverage without the harness caveat attached. The like-for-like comparison, 62.7 percent against 7.8 percent, is still a large jump. It is also the only version of that comparison the benchmark actually supports.

Then the published figures themselves shifted. According to The Next Web’s account of Fortune’s reporting, journalist Emily Forlini lined up archived versions of OpenAI’s launch post and turned up five separate metrics that had been changed since the post first went live. Astra’s hallucination rate moved from 4.2 percent to 2 percent and back to 4.2 percent. Anthropic’s Fable 5.1 score on FrontierMath dropped from 87.8 percent to 78 percent before settling at 83 percent. On ExploitBench, Sol’s listed score jumped from 5.5 percent to 11.5 percent, and OpenAI confirmed to Fortune that it is now looking into rolling that change back. At one point OpenAI took the entire launch post offline and then put it back up, declining to say why.

At Stanford, researchers Anka Reuel and Mike Hardy of the Intelligent Systems Laboratory and the Trustworthy AI Lab coined a term for the pattern: benchmaxxing, in which an evaluation gets rerun again and again under shifting conditions until the score finally climbs. They also examined Astra’s system card and found what they described as barely any documentation of the internal hallucination benchmark’s methodology, down to how many items it even contains. Not every source reads this as deliberate. Snorkel AI’s Vincent Sunn Chen gave Fortune a gentler read, noting that checkpoints, configurations, and grading criteria are all still shifting right up until launch, which is why last-minute score movement is common; his proposed fix is a norm obligating companies to disclose exactly what changed whenever they revise a published number.

None of this establishes that OpenAI acted in bad faith. What the record shows is a company that built its central capability claim on the version of a benchmark most favorable to it, then revised several of the supporting metrics after publication without initially flagging the changes. Any operator citing Astra’s ARC-AGI-3 score in a procurement decision or a competitive analysis should now cite the harness alongside the number, and should ask every vendor claiming a benchmark result which scaffold produced it before treating that figure as comparable to anyone else’s.

The Next Web (Ana Maria Constantin), published September 6, 2026.