BrowseComp
Often citedHigher is better
OpenAI's web-browsing test: 1,266 hard questions whose answers can only be found by searching repeatedly and piecing clues together. Answers are short and easy to check but hard to find. Scores run from 0 to 100%. Higher is better.
This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.
Top score92.2%GPT-5.6 Sol
Models tested40
Last updated2026.09.20
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives answers people prefer most, without false claims?Arena Text Factuality
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it know facts accurately without searching?SimpleQA Verified
GPT-6 Astra🥇
GPT-6 Astra🥇Can it carry out multi-step web research?Deep Research Bench
Claude Opus 4.6🥇
Claude Opus 4.6🥇
- 🥇OpenAIGPT-5.6 SolReasoning effort: Max92.2%
- 🥈OpenAIGPT-6 Astra91.5%
- 🥉Moonshot AIKimi K3Reasoning effort: Max91.2%
- #4AnthropicClaude Opus 5Reasoning effort: Max90.8%
- #5OpenAIGPT-5.4 ProReasoning effort: Extra High89.3%
- #6StepfunStep-5Reasoning effort: High88.7%
- #7AnthropicClaude Opus 4.8Reasoning effort: Max88.5%
- #8OpenAIGPT-5.6 Terra87.5%
- #9AnthropicClaude Fable 587.4%
- #10AnthropicClaude Sonnet 5Reasoning effort: Max86.6%
- #11AnthropicClaude Opus 4.6Reasoning effort: Max86.6%
- #12Moonshot AIKimi K2.6Reasoning effort: Thinking86.3%
- #13GoogleGemini 3.1 ProReasoning effort: High85.9%
- #14OpenAIGPT-5.5Reasoning effort: Extra High84.4%
- #15MiniMaxMiniMax M383.5%
- #16DeepSeekDeepSeek V4 ProReasoning effort: Max83.4%
- #17OpenAIGPT-5.6 Luna83.3%
- #18OpenAIGPT-5.4Reasoning effort: Extra High82.7%
- #19AnthropicClaude Sonnet 4.6Reasoning effort: Max82.1%
- #20MeituanLongCat 2.079.9%
- #21AnthropicClaude Opus 4.779.8%
- #22Moonshot AIKimi K2.5Reasoning effort: Thinking78.4%
- #23ByteDanceDola Seed 2.0 Pro77.3%
- #24MiniMaxMiniMax M2.776.3%
- #24MiniMaxMiniMax M2.576.3%
- #26StepfunStep 3.7 Flash75.8%
- #27DeepSeekDeepSeek V4 FlashReasoning effort: Max73.2%
- #28Z.aiGLM-5.168.0%
- #29AnthropicClaude Opus 4.5Reasoning effort: Max (no thinking)67.8%
- #30TencentHy367.1%
- #31Z.aiGLM-562.0%
- #32DeepSeekDeepSeek V3.2Reasoning effort: Thinking51.4%
- #33UpstageSolar Pro 449.2%
- #34Mistral AIMistral Medium 3.5Reasoning effort: Max48.6%
- #35NVIDIANemotron 3 Ultra44.4%
- #36AnthropicClaude Sonnet 4.543.9%
- #37LG AI ResearchK-EXAONEReasoning effort: None31.4%
- #38NVIDIANemotron 3 Super31.3%
- #39Mistral AIMistral Small 421.3%
- #40SK TelecomA.X K2Reasoning effort: Thinking9.3%
Half of models ≤ 79.8%Measured by: Maker-reportedLast updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.
Vendor-reported numbers use varying setups — treat them as indicative.