Vectara Factual Consistency

In plain words · How often its summaries stick to what the document actually says

What does it measure?
How faithfully a document summary sticks to the source. It is close to the hallucination rate flipped around (roughly 100% − hallucination rate) — the same information viewed as 'higher is better'.
Who evaluates it, and how?
The search company Vectara has AIs summarise more than 7,700 documents spanning news, technology, medicine, law and more, and its own judge model HHEM (currently version 2.3) checks whether the summary contains facts absent from the source. The document set is private, so it cannot be learned in advance. This table takes the values from the public leaderboard; an AI that has since been removed or renamed there is listed separately below the table with its last published value, unranked.
Reading the score
0–100%, higher is better. It nearly mirrors the hallucination rate, but the two are published separately and can differ slightly in the decimals. It measures summarisation faithfulness only and says nothing about general knowledge accuracy.

Last updated: 2026-09-22

Rank

Top score96.9%GPT-5.4 Nano
Models tested36
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.03.17
    96.9%
  2. 2
    Google2025.09.25
    96.7%
  3. 3
    OpenAI2026.03.17
    94.5%
  4. 4
    OpenAI2025.04.14
    94.4%
  5. 5
    DeepSeek2025.12.01
    93.7%
  6. 6
    OpenAI2026.09.22
    93.5%
  7. 7
    Arcee AI2026.04.01
    93.1%
  8. 8
    Google2025.06.17
    93.0%
  9. 9
    OpenAI2026.03.05
    93.0%
  10. 10
    Google2026.04.02
    92.6%

35 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 32 more

Source: Vectara HHEM Data sources & removal requests

What to keep in mind

  • This score measures whether a summary stays faithful to its source document — not general knowledge or reasoning.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.