Bench to the Future 3
Lower is better
A forecasting test by FutureSearch: agents forecast real questions that have already been resolved, researching only a frozen past copy of the web so the answers cannot be looked up. It mixes yes/no and numeric questions and is scored on forecasting error (Brier score). Scores are forecasting error (Brier). Lower is better.
Top score0.120Claude Opus 5
Models tested10
Model release date (newest on the right)
Best score so farLower scores are better here, so the axis is flipped: higher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1FutureSearch agent · xhigh2026.07.24$200.120
- 2FutureSearch agent · high2026.06.09$400.130
- 3FutureSearch agent · xhigh2026.05.27$200.132
- 4Vendor agent SDK · high2026.04.23$240.134
- 5FutureSearch agent · low2026.09.04$400.135
- 5FutureSearch agent · high2026.07.09$160.135
- 7FutureSearch agent · xhigh2026.09.02$3.500.141
- 8FutureSearch agent · xhigh2026.06.30$80.142
- 9FutureSearch agent2026.08.18$3.600.149
- 9FutureSearch agent2026.08.26$0.410.149
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.