Vectara Hallucination Rate

In plain words · How often it makes things up when summarizing a document

What does it measure?
How often the AI invents content that is not in the source when asked to summarise a document. Lower means more faithful to the source. It is about 'staying within the given text', not 'knowing a lot'.
Who evaluates it, and how?
The search company Vectara has AIs summarise more than 7,700 documents spanning news, technology, medicine, law and more, and its own judge model HHEM (currently version 2.3) checks whether the summary contains facts absent from the source. The document set is private, so it cannot be learned in advance. This table takes the values from the public leaderboard; an AI that has since been removed or renamed there is listed separately below the table with its last published value, unranked.
Reading the score
0–100%, lower is better. The top models cluster in the 1–3% range, so decimal differences mean little. It measures faithfulness in summarisation only, so it says nothing about fabrication in general questions or about summary quality and usefulness.

Last updated: 2026-09-22

Rank

Top score3.1%GPT-5.4 Nano
Models tested36
Last updated2026.09.22
Model release date (newest on the right)
Best score so farLower scores are better here, so the axis is flipped: higher on the chart is better
  1. 🥇OpenAIGPT-5.4 Nano3.1%
  2. 🥈GoogleGemini 2.5 Flash Lite3.3%
  3. 🥉OpenAIGPT-5.4 Mini5.5%
  4. #4OpenAIGPT-4.15.6%
  5. #5DeepSeekDeepSeek V3.26.3%
  6. #6OpenAIGPT-6 Sol6.5%
  7. #7Arcee AITrinity Large Thinking6.9%
  8. #8GoogleGemini 2.5 Pro7.0%
  9. #8OpenAIGPT-5.47.0%
  10. #10GoogleGemma 4 31B7.4%
  11. #11GoogleGemini 2.5 Flash7.8%
  12. #12GoogleGemini 3.1 Flash Lite8.2%
  13. #13OpenAIGPT-5.4 Pro8.3%
  14. #14DeepSeekDeepSeek V4 Pro8.6%
  15. #15OpenAIGPT-6 Astra8.7%
  16. #16OpenAIGPT-5.59.3%
  17. #17NVIDIANemotron 3 Nano 30B A3B9.6%
  18. #18AnthropicClaude Haiku 4.59.8%
  19. #19Z.aiGLM-510.1%
  20. #20AnthropicClaude Sonnet 410.3%
  21. #21GoogleGemini 3.1 Pro10.4%
  22. #22OpenAIGPT-5 Nano10.5%
  23. #23AnthropicClaude Sonnet 4.610.6%
  24. #24AlibabaQwen3.5 Plus10.7%
  25. #25AnthropicClaude Opus 4.510.9%
  26. #26AnthropicClaude Opus 4.111.8%
  27. #27AnthropicClaude Opus 412.0%
  28. #27AnthropicClaude Opus 4.712.0%
  29. #27AnthropicClaude Sonnet 4.512.0%
  30. #30AnthropicClaude Opus 4.612.2%
  31. #31OpenAIGPT-5.6 Sol12.4%
  32. #32OpenAIGPT-5 Mini12.9%
  33. #33GoogleGemini 3 Flash13.5%
  34. #34OpenAIGPT OSS 120B14.2%
  35. #35xAIGrok 4.1 Fast17.8%
  36. #36xAIGrok 4.1 Fast (Reasoning)19.2%
Half of models ≤ 10.0%Last updated 2026-09-22

Scores from an earlier edition

These scores are not in the latest leaderboard (2026.09.22) — the model was dropped by the source or has not been refreshed yet. The last score we received is shown for reference, unranked.

  1. —
    Meta2025.04.05
    7.7%
  2. —
    Meta2025.04.05
    8.2%
  3. —
    Z.ai2026.04.07
    10.1%
  4. —
    Anthropic2026.05.27
    12.0%
  5. —
    Moonshot AI2026.01.27
    14.2%
  6. —
    OpenAI2025.08.07
    14.7%

32 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Perplexity Decider v1 27B, Clef Flash and 29 more

Source: Vectara HHEM Data sources & removal requests

What to keep in mind

  • This score measures whether a summary stays faithful to its source document — not general knowledge or reasoning.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.