MRCR v2 (8-needle)
Widely citedHigher is better
A long-context test: a very long synthetic conversation hides eight nearly identical requests (such as several poems on the same topic), and the model must reproduce one specific instance, such as 'the third one'. This is Google DeepMind's MRCR v2; models are tested up to different maximum lengths, so their averages cover different ranges. Scores run from 0 to 100%. Higher is better.
Top score88.1%Qwen3.6 Max
Models tested53
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1enabled2026.04.27$4.8888.1%
- 2max2026.07.09$1685.5%
- 3enabled2026.04.27$0.7883.5%
- 4high2026.08.13$383.0%

- 5xhigh2026.04.23$2479.8%
- 6max2026.07.24$2077.3%
- 7max2026.07.09$9.5076.6%
- 8enabled2026.02.16$2.8075.5%
- 9medium2026.08.12$573.6%
- 10high2026.07.21$373.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.