MirrorCode

Higher is better

Made by Epoch AI with METR: the AI reimplements an existing program (Unix tools, compression, cryptography, interpreters, and more) from scratch without seeing its source code. It succeeds only if every output matches the original exactly, including on hidden tests, with up to seven days per attempt. Scores run from 0 to 100%. Higher is better.

Top score77.4%Claude Opus 5.5
Models tested9
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.09.22 · max
    77.4%
  2. 2
    Anthropic2026.09.01 · high
    73.3%
  3. 3
    Anthropic2026.06.09 · high
    63.9%
  4. 4
    OpenAI2026.09.04 · high
    46.7%
  5. 5
    Anthropic2026.04.16 · high
    31.1%
  6. 6
    OpenAI2026.07.09 · high
    20.0%
  7. 7
    OpenAI2026.03.05 · high
    15.6%
  8. 8
    OpenAI2026.04.23 · high
    10.0%
  9. 9
    Google2026.02.19 · high
    8.9%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.