Fiction.LiveBench (16k)
Often citedHigher is better
A long-text comprehension test on stories from the fiction site Fiction.live: questions require inferring characters' thoughts, the order of events, and unstated information. The value here is the score when the story is about 16k tokens long; scores differ greatly at other lengths. Scores run from 0 to 100%. Higher is better.
Top score97.2%GPT-5
Models tested10
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it pull scattered facts from several long documents into one answer?Long Context Reasoning
Kimi K3🥇
Kimi K3🥇Can it read tables and reshape them as needed?LiveBench Data Analysis
GPT-6 Astra🥇
GPT-6 Astra🥇Which AI gives the business answers people prefer most?Arena Business, Management & Finance
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇OpenAIGPT-5Reasoning effort: Medium97.2%
- 🥈Moonshot AIKimi K2.586.1%
- 🥉OpenAIGPT-5 MiniReasoning effort: Medium69.4%
- #4OpenAIGPT-4.163.9%
- #5AnthropicClaude Opus 461.1%
- #6AnthropicClaude Sonnet 446.9%
- #7MetaLlama 4 Maverick46.2%
- #8OpenAIGPT OSS 120BReasoning effort: High44.4%
- #8OpenAIGPT-5 NanoReasoning effort: Medium44.4%
- #10MetaLlama 4 Scout36.0%
#1 leads #2 by 11.1 pts · Half of models ≤ 54.0%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.