FACTS Grounding

Higher is better

Given a document up to 32k tokens long and a request (summarize, find facts, analyze), the model writes a long answer grounded only in that document. AI judges check whether it addresses the request and whether every claim is supported by the document. Scores run from 0 to 100%. Higher is better.

Top score85.2%GPT-6 Sol
Models tested31
Last updated2026.09.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.09.22
    85.2%
  2. 2
    OpenAI2026.09.04
    82.0%
  3. 3
    OpenAI2026.09.22
    81.4%
  4. 4
    Google2026.04.02
    80.7%
  5. 5
    OpenAI2026.04.23
    76.2%
  6. 6
    Google2026.08.13
    75.5%
  7. 7
    xAI2026.07.08
    74.5%
  8. 8
    Google2025.06.17
    74.3%
  9. 9
    Google2026.05.19
    73.2%
  10. 10
    Google2026.09.03
    73.1%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.