FACTS Multimodal
Higher is better · Image-input models only
Answer questions about images factually, combining what is in the image with general knowledge. An answer counts only if it includes the essential facts and contradicts nothing. Only models that accept images can take part. Scores run from 0 to 100%. Higher is better.
Top score50.7%Gemini 3.8 Flash
Models tested31
Last updated2026.09.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12026.09.03$350.7%

- 22026.08.13$349.8%

- 32026.09.22$1648.4%
- 42026.09.04$4047.7%
- 52026.07.21$347.2%

- 62025.06.17$7.8146.9%

- 72026.04.23$2445.8%
- 82025.08.07$7.8144.1%
- 92026.09.22$843.1%
- 102025.08.07$1.5641.5%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.