Fiction.LiveBench (16k)

Often citedHigher is better

A long-text comprehension test on stories from the fiction site Fiction.live: questions require inferring characters' thoughts, the order of events, and unstated information. The value here is the score when the story is about 16k tokens long; scores differ greatly at other lengths. Scores run from 0 to 100%. Higher is better.

Top score97.2%GPT-5
Models tested10
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇OpenAIGPT-5Reasoning effort: Medium97.2%
  2. 🥈Moonshot AIKimi K2.586.1%
  3. 🥉OpenAIGPT-5 MiniReasoning effort: Medium69.4%
  4. #4OpenAIGPT-4.163.9%
  5. #5AnthropicClaude Opus 461.1%
  6. #6AnthropicClaude Sonnet 446.9%
  7. #7MetaLlama 4 Maverick46.2%
  8. #8OpenAIGPT OSS 120BReasoning effort: High44.4%
  9. #8OpenAIGPT-5 NanoReasoning effort: Medium44.4%
  10. #10MetaLlama 4 Scout36.0%
#1 leads #2 by 11.1 pts · Half of models ≤ 54.0%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.