FACTS Benchmark Suite

Often citedHigher is better

The overall score of Google's FACTS factuality suite: the average of four parts, namely answering from memory (Parametric), answering with a search tool (Search), questions about images (Multimodal), and long answers grounded only in a given document (Grounding). Models not tested on all four parts are hard to compare. Scores run from 0 to 100%. Higher is better.

Top score75.7%GPT-6 Astra
Models tested31
Last updated2026.09.29
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.09.04
    75.7%
  2. 2
    OpenAI2026.09.22
    72.5%
  3. 3
    Google2026.08.13
    71.2%
  4. 4
    Google2026.09.03
    69.4%
  5. 5
    Google2026.02.19
    67.4%
  6. 6
    OpenAI2026.04.23
    66.7%
  7. 7
    Google2026.07.21
    65.3%
  8. 8
    Anthropic2026.09.22
    65.3%
  9. 9
    Google2025.12.17
    63.4%
  10. 10
    OpenAI2026.09.22
    63.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.