The Forecasting Research Institute, a nonprofit that has run expert and “superforecaster” panels on AI progress since 2022, published a review this week grading how those predictions held up. The headline result: on benchmark scores specifically, forecasters have consistently guessed low, sometimes by years rather than months.

On FrontierMath, a math benchmark, the median expert forecast for the best model’s end of 2025 score was 31 percent, with superforecasters at 30 percent. The benchmark actually resolved at 40.7 percent. A coding benchmark called LiveCodeBench Pro (Hard) shows a wider gap: experts predicted 14 percent and superforecasters 12 percent for the state of the art by the end of 2026, but the leading model had already reached 53.8 percent by May 2026, more than three times the forecast.

This is not a new pattern. Back in a 2022 tournament run before ChatGPT’s public release, the outcomes that eventually happened looked like long shots to both panels: superforecasters put just a 9.7 percent chance on them, domain experts 24.6 percent. Neither group treated the coming reality as likely. The starkest single miss: forecasters expected an AI system to reach International Mathematical Olympiad gold medal performance around 2030 (experts) or 2035 (superforecasters). A model cleared that bar in July 2025.

Revenue forecasts followed the same shape. In one FRI study, AI industry experts put a leading AI company’s annual recurring revenue at $20 billion by the end of 2026, economists guessed $16 billion, and superforecasters guessed $25 billion. The institute says the actual figure, as of September 22, 2026, is likely close to $100 billion. A separate panel asked forecasters to predict Anthropic and OpenAI’s combined annualized revenue run rate by year end: the median guesses were $70 billion (experts) and $90 billion (superforecasters), against a probable current figure near $140 billion.

Adoption and diffusion forecasts come out more mixed, though the institute says they still lean toward underestimation overall. One 2025 question asked how much of US work time generative AI was assisting: experts guessed 4 percent, superforecasters guessed 3.6 percent, and the figure actually resolved at 5.7 percent. FRI notes the forecasters may have anchored to a baseline number they were given that was later revised upward. Not every question broke that way: experts overestimated how much of the US ride-hailing market autonomous vehicles would capture by 2027, predicting 7.3 percent against a roughly 2.5 percent trend, while superforecasters landed closer at 2 percent. Predictions tied to physical infrastructure fared better: data center buildout and how much power AI would draw from the US grid both tracked much closer to reality than revenue or usage figures did.

On the biggest question, whether AI is reshaping employment, GDP growth, or life expectancy at a macro scale, the institute says there simply is not enough resolved data yet to score anyone. Most of those questions run years past their target dates. FRI is also upfront that its own method is skewed toward catching underestimates rather than overestimates: a forecast can be proven too low the moment reality passes it, but proving one too high requires waiting out the full deadline, which most of these questions have not reached.

For operators building roadmaps or fundraising decks around analyst timelines for model capability, the practical lesson is to widen the error bars upward rather than split the difference. A planning horizon built on the median expert view of when a capability arrives has, across four years of FRI’s own data, been the version most likely to be broken by a launch that arrived early.

Forecasting Research Institute published this analysis on September 23, 2026, based on forecasts collected since 2022 across its Existential Risk Persuasion Tournament and Longitudinal Expert AI Panel studies.