The Forecasting Research Institute (FRI) found that top AI experts and economists significantly underestimated recent progress in AI capabilities. AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years before the median expert forecast and ten years before the median superforecaster forecast.

Experts assigned an average probability of 24.6 percent to the benchmark results that actually happened, while superforecasters assigned just 9.7 percent. For gold-medal performance at the Math Olympiad, the figures dropped to 8.6 and 2.3 percent. The predictions were gathered in 2022, before ChatGPT launched, but the pattern held afterward too, according to FRI.

AI may also have solved a Millennium Prize Problem, though it's still unclear whether the solution meets the evaluation criteria.

In a survey from August and September 2025, experts had put the median odds of such a solution by the end of 2027 at just 10 percent, and superforecasters at 5.4 percent.

In a study of AI capabilities in virology, experts predicted AI models wouldn't match a top team of virologists on a troubleshooting benchmark until 2030. Superforecasters said 2034. FRI says that likely happened as early as April 2025.

A cybersecurity benchmark showed similar underestimates. At the median, experts expected AI to match a top team on the Virology Capabilities Test by 2030, and superforecasters by 2034. FRI says it likely happened in April 2025.

Respondents tied this milestone to higher expected biorisk, not to any documented rise in actual harm.

Economic forecasts were also far too conservative. Experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion. Economists said $16 billion, and superforecasters said $25 billion.

FRI cites roughly $100 billion for Anthropic in September 2026 as a figure that has likely already been reached. Annualized revenue at Anthropic and OpenAI far exceeded the median forecasts of every surveyed group.

Real-world impact is harder to call. Not every forecast ran too low. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks.

Virologists expected 40 percent, superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone. The language model made no measurable difference, though the trial was small.

Source: thedecoder