PostTrainBench

Often citedHigher is better

Made by researchers in Tübingen: the agent gets four small base language models, one GPU, and 10 hours per run, and must post-train them to raise their scores on seven tests covering math, code, medicine, and more. The result is the weighted average of the trained models' scores. Scores run from 0 to 100%. Higher is better.

Top score41.8%Claude Fable 5
Models tested11
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.06.09 · Claude Code · max
    41.8%
  2. 2
    OpenAI2026.07.09 · Codex CLI · max
    36.2%
  3. 3
    Anthropic2026.07.24 · Claude Code
    35.0%
  4. 4
    Anthropic2026.05.27 · Claude Code · high
    33.8%
  5. 5
    Moonshot AI2026.07.17 · Claude Code
    32.0%
  6. 6
    Z.ai2026.06.16 · Claude Code · max
    31.7%
  7. 7
    Anthropic2026.04.16 · Claude Code · xhigh
    28.6%
  8. 8
    OpenAI2026.04.23 · Codex CLI · xhigh
    27.2%
  9. 9
    xAI2026.07.08 · Cursor CLI · high
    23.4%
  10. 10
    Google2026.02.19 · OpenCode
    22.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.