FACTS Benchmark Suite
Often citedHigher is better
The overall score of Google's FACTS factuality suite: the average of four parts, namely answering from memory (Parametric), answering with a search tool (Search), questions about images (Multimodal), and long answers grounded only in a given document (Grounding). Models not tested on all four parts are hard to compare. Scores run from 0 to 100%. Higher is better.
Top score75.7%GPT-6 Astra
Models tested31
Last updated2026.09.29
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12026.09.04$4075.7%
- 22026.09.22$872.5%
- 32026.08.13$371.2%

- 42026.09.03$369.4%

- 52026.02.19$9.5067.4%

- 62026.04.23$2466.7%
- 72026.07.21$365.3%

- 82026.09.22$1665.3%
- 92025.12.17$2.3863.4%

- 102026.09.22$0.4063.3%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.