Deep Research Bench (FutureSearch)
Higher is better
A web-research test by FutureSearch: the agent searches a frozen offline copy of the web to do multi-step research tasks such as finding a number, compiling data, or checking a claim. Each task is scored 0 to 1 and averaged. It is a different test from the similarly named 'DeepResearch Bench'. Scores run from 0 to 100%. Higher is better.
Top score55.3%Claude Opus 4.6
Models tested17
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives answers people prefer most, without false claims?Arena Text Factuality
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it know facts accurately without searching?SimpleQA Verified
GPT-6 Astra🥇
GPT-6 Astra🥇
- 🥇AnthropicClaude Opus 4.6Reasoning effort: High149447.0%55.3%
- 🥈AnthropicClaude Sonnet 4.6Reasoning effort: High146135.5%54.9%
- 🥉AnthropicClaude Opus 4.5Reasoning effort: High145945.7%54.8%
- #4OpenAIGPT-5.5Reasoning effort: High148663.0%54.0%
- #5AnthropicClaude Sonnet 4.5Reasoning effort: Budget 2K144930.7%52.6%
- #6AnthropicClaude Opus 4.8Reasoning effort: High147053.0%50.2%
- #7GoogleGemini 3 FlashReasoning effort: Low146866.8%49.8%
- #8OpenAIGPT-5Reasoning effort: Low142150.1%49.6%
- #9AnthropicClaude Opus 4.11435—48.3%
- #10GoogleGemini 3.1 ProReasoning effort: High147273.5%47.8%
- #11AnthropicClaude Opus 4Reasoning effort: Budget 2K——46.8%
- #12AnthropicClaude Sonnet 4Reasoning effort: Budget 2K——46.6%
- #13AnthropicClaude Haiku 4.5Reasoning effort: Low141813.2%45.5%
- #14GoogleGemini 2.5 Pro1462—41.5%
- #15GoogleGemini 3.1 Flash LiteReasoning effort: Low1433—37.3%
- #16OpenAIGPT-5.4 MiniReasoning effort: Low144729.4%36.3%
- #17OpenAIGPT-5.4Reasoning effort: Low148545.1%35.1%
Half of models ≤ 48.3%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.