FACTS Grounding
Higher is better
Given a document up to 32k tokens long and a request (summarize, find facts, analyze), the model writes a long answer grounded only in that document. AI judges check whether it addresses the request and whether every claim is supported by the document. Scores run from 0 to 100%. Higher is better.
Top score85.2%GPT-6 Sol
Models tested31
Last updated2026.09.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12026.09.22$885.2%
- 22026.09.04$4082.0%
- 32026.09.22$0.4081.4%
- 42026.04.02$0.3480.7%

- 52026.04.23$2476.2%
- 62026.08.13$375.5%

- 72026.07.08$574.5%
- 82025.06.17$7.8174.3%

- 92026.05.19$7.1373.2%

- 102026.09.03$373.1%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.