A test OpenAI built less than a year ago to check whether AI could do real office work now reads like a relic, according to an essay that circulated widely on X this week. The author, who posts under the handle knowerofmarkets, argues that the pace of change since that test launched is the clearest evidence yet that frontier AI is closing in on something like superintelligence.

The test in question is GDPval, an OpenAI evaluation from roughly a year ago that graded AI models against human professionals with 10 to 15 years of experience across 44 jobs and nine industries, from compliance officers to industrial engineers. In OpenAI’s own September 2025 results, Claude Opus 4.1 scored 38.8 percent and GPT-5-High scored 47.6 percent against those human benchmarks. Those are OpenAI’s self-reported figures, not an independent audit.

The essay’s central move is a before-and-after comparison. GPT-5’s release blog touted a score of 13.5 percent on FrontierMath and 61.9 percent on AIME 2025 without extended reasoning. The author says those same benchmarks barely register in more recent model announcements, replaced by newer tests such as Terminal-Bench 4.0, GeneBench Pro, OSWorld 2.0, and Automation Bench.

The more contested claim in the piece involves mathematics. The author cites a September 8 announcement that an unreleased OpenAI model solved the Navier-Stokes Millennium Prize Problem, and relays a secondhand claim, sourced to set theorist Elliot Glazer speaking on a podcast, that OpenAI may be sitting on roughly 100 solutions to other longstanding math problems. That claim has not been independently verified, and it drew a public rebuke: a group of mathematicians published an open letter arguing that “the goals of the AI companies and the goals of the mathematical community are severely misaligned.”

On the money side, the essay cites Axios reporting that Anthropic’s annualized revenue is now pacing toward $100 billion, up from estimates of $50 billion to $60 billion just a couple of months earlier. That is a steep jump for any company to report in that window, and it is worth remembering that annualized run rates can swing hard in either direction if a handful of large contracts land or lapse.

The essay treats benchmark churn itself as proof of an exponential curve. That skips an obvious alternative explanation: benchmarks get replaced precisely when models start saturating them, which is what a working evaluation pipeline is supposed to do, not evidence on its own that the underlying systems have entered a new category. A test going stale in a year says as much about how fast labs iterate their scorecards as it does about the models being scored.

For operators, the more useful question is not whether any of this adds up to superintelligence. It is whether claims like an unreleased model solving a Millennium Prize Problem get independently checked before they show up in the next funding pitch or the next earnings comparison. Anyone using Anthropic’s reported revenue trajectory to model a competitor’s runway should treat the $100 billion figure as a pacing estimate from a single outlet, not an audited number, until the company confirms it directly.

Reported in an essay the author published on X (@knowerofmarkets) on 24 September 2026, citing Axios reporting from 18 September 2026.