Bench to the Future 3

Lower is better

A forecasting test by FutureSearch: agents forecast real questions that have already been resolved, researching only a frozen past copy of the web so the answers cannot be looked up. It mixes yes/no and numeric questions and is scored on forecasting error (Brier score). Scores are forecasting error (Brier). Lower is better.

Top score0.120Claude Opus 5
Models tested10
Model release date (newest on the right)
Best score so farLower scores are better here, so the axis is flipped: higher on the chart is better
  1. 1
    Anthropic2026.07.24 · FutureSearch agent · xhigh
    0.120
  2. 2
    Anthropic2026.06.09 · FutureSearch agent · high
    0.130
  3. 3
    Anthropic2026.05.27 · FutureSearch agent · xhigh
    0.132
  4. 4
    OpenAI2026.04.23 · Vendor agent SDK · high
    0.134
  5. 5
    OpenAI2026.09.04 · FutureSearch agent · low
    0.135
  6. 5
    OpenAI2026.07.09 · FutureSearch agent · high
    0.135
  7. 7
    Meta2026.09.02 · FutureSearch agent · xhigh
    0.141
  8. 8
    Anthropic2026.06.30 · FutureSearch agent · xhigh
    0.142
  9. 9
    Z.ai2026.08.18 · FutureSearch agent
    0.149
  10. 9
    Z.ai2026.08.26 · FutureSearch agent
    0.149

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.