Ko-TruthfulQA
KoreanHigher is better
TruthfulQA-style questions designed to draw out common misconceptions, translated into Korean and answered as multiple choice. It checks whether the model resists plausible but wrong answers. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score98.0%Claude Fable 5
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.06.09$4098.0%
- 1high-effort2026.02.19$9.5098.0%

- 3xhigh-effort2026.04.23$2496.0%
- 3high-effort2026.05.27$2096.0%
- 3high-effort2025.05.22$6096.0%
- 3high-effort2025.09.29$1296.0%
- 7high-effort2026.06.30$895.0%
- 7high-effort2025.12.17$2.3895.0%

- 7high-effort2025.11.24$2095.0%
- 72026.05.21$3.6995.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.