MRCR v2 (8-needle)

Widely citedHigher is better

A long-context test: a very long synthetic conversation hides eight nearly identical requests (such as several poems on the same topic), and the model must reproduce one specific instance, such as 'the third one'. This is Google DeepMind's MRCR v2; models are tested up to different maximum lengths, so their averages cover different ranges. Scores run from 0 to 100%. Higher is better.

Top score88.1%Qwen3.6 Max
Models tested53
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Alibaba2026.04.27 · enabled
    88.1%
  2. 2
    OpenAI2026.07.09 · max
    85.5%
  3. 3
    Alibaba2026.04.27 · enabled
    83.5%
  4. 4
    Google2026.08.13 · high
    83.0%
  5. 5
    OpenAI2026.04.23 · xhigh
    79.8%
  6. 6
    Anthropic2026.07.24 · max
    77.3%
  7. 7
    OpenAI2026.07.09 · max
    76.6%
  8. 8
    Alibaba2026.02.16 · enabled
    75.5%
  9. 9
    xAI2026.08.12 · medium
    73.6%
  10. 10
    Google2026.07.21 · high
    73.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.