A new preprint argues that frontier language models are far better at physics than the leaderboards suggest, because the leaderboards themselves are broken. Researchers led by Ali Ansari re-graded six widely used physics benchmarks with faculty and graduate experts and found that most answers marked wrong were not model errors at all.

The culprits were mundane: incorrect reference solutions, grading scripts that failed to parse a valid answer, and questions that were ambiguous or underspecified in the first place. Once those flaws were fixed or the affected questions excluded, one frontier model’s scores jumped sharply across the retested benchmarks.

The paper reports GPT-5.6-Sol’s mean@4 score rising from 47.3% to 78.7% on HLE-Physics, and from 61.0% to 87.2% on the CMT-Benchmark, after expert correction. On the 54 retained CritPt challenges, its corrected pass@4 reached 94.4%. The researchers say scores on the audited portions of UGPhysics, PRISM-Physics, and PHYBench rose substantially as well, though the paper does not give exact figures for those three in the abstract.

This is the researchers’ own reported result from a single preprint, posted to arXiv and not yet peer reviewed. It has not been independently replicated, and the corrected numbers apply only to the retained, expert-reviewed subsets of each benchmark, not the full original test sets.

The stakes go beyond one model’s report card. The benchmarks named here, including ones featured in the Artificial Analysis Intelligence Index, sit inside procurement decisions, model comparisons, and press coverage that treats a raw percentage as a settled measurement of capability. If a meaningful share of “failures” on those tests trace back to bad answer keys rather than bad reasoning, then every past comparison built on the uncorrected scores was measuring benchmark quality as much as model quality.

That is the sharper version of a problem builders already suspect: a model that scores low on a niche eval might simply be getting graded against a wrong answer. The paper’s own framing supports this. It states plainly that experts who rely on these models day to day often do not see the difficulties the raw scores suggest, a gap between reported failure and observed competence that should make anyone citing a benchmark score pause before treating it as ground truth.

The more durable finding may be the near-saturation itself. If several of the field’s toughest closed-ended physics tests are already close to maxed out once cleaned up, they stop being useful as a way to separate frontier models from each other. The paper’s own conclusion points the same direction: what is needed next is harder, expert-validated evaluation, not another leaderboard built on a quiz nobody proofread.

For any team benchmarking models on specialized domains, from physics to law to biology, this is a reason to spot-check a sample of the questions a model gets marked wrong on before concluding the model is weak. The answer key might be the thing that needs fixing.

Ansari, Ali, et al. “How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks.” arXiv preprint, posted September 11, 2026.