PostTrainBench
Often citedHigher is better
Made by researchers in Tübingen: the agent gets four small base language models, one GPU, and 10 hours per run, and must post-train them to raise their scores on seven tests covering math, code, medicine, and more. The result is the weighted average of the trained models' scores. Scores run from 0 to 100%. Higher is better.
Top score41.8%Claude Fable 5
Models tested11
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1Claude Code · max2026.06.09$4041.8%
- 2Codex CLI · max2026.07.09$1636.2%
- 3Claude Code2026.07.24$2035.0%
- 4Claude Code · high2026.05.27$2033.8%
- 5Claude Code2026.07.17$1132.0%
- 6Claude Code · max2026.06.16$3.5731.7%
- 7Claude Code · xhigh2026.04.16$2028.6%
- 8Codex CLI · xhigh2026.04.23$2427.2%
- 9Cursor CLI · high2026.07.08$523.4%
- 10OpenCode2026.02.19$9.5022.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.