Deep Research Bench (FutureSearch)

Higher is better

A web-research test by FutureSearch: the agent searches a frozen offline copy of the web to do multi-step research tasks such as finding a number, compiling data, or checking a claim. Each task is scored 0 to 1 and averaged. It is a different test from the similarly named 'DeepResearch Bench'. Scores run from 0 to 100%. Higher is better.

Top score55.3%Claude Opus 4.6
Models tested17
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 4.6Reasoning effort: High149447.0%55.3%
  2. 🥈AnthropicClaude Sonnet 4.6Reasoning effort: High146135.5%54.9%
  3. 🥉AnthropicClaude Opus 4.5Reasoning effort: High145945.7%54.8%
  4. #4OpenAIGPT-5.5Reasoning effort: High148663.0%54.0%
  5. #5AnthropicClaude Sonnet 4.5Reasoning effort: Budget 2K144930.7%52.6%
  6. #6AnthropicClaude Opus 4.8Reasoning effort: High147053.0%50.2%
  7. #7GoogleGemini 3 FlashReasoning effort: Low146866.8%49.8%
  8. #8OpenAIGPT-5Reasoning effort: Low142150.1%49.6%
  9. #9AnthropicClaude Opus 4.11435—48.3%
  10. #10GoogleGemini 3.1 ProReasoning effort: High147273.5%47.8%
  11. #11AnthropicClaude Opus 4Reasoning effort: Budget 2K——46.8%
  12. #12AnthropicClaude Sonnet 4Reasoning effort: Budget 2K——46.6%
  13. #13AnthropicClaude Haiku 4.5Reasoning effort: Low141813.2%45.5%
  14. #14GoogleGemini 2.5 Pro1462—41.5%
  15. #15GoogleGemini 3.1 Flash LiteReasoning effort: Low1433—37.3%
  16. #16OpenAIGPT-5.4 MiniReasoning effort: Low144729.4%36.3%
  17. #17OpenAIGPT-5.4Reasoning effort: Low148545.1%35.1%
Half of models ≤ 48.3%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.