Output Speed (tokens/s)

In plain words · How fast it writes — chunks of text produced per second

What does it measure?
How fast the AI writes its answer. For an answer of the same length, more tokens (roughly, chunks of characters) per second means less waiting.
Who evaluates it, and how?
Output speed measured per provider on real requests routed through OpenRouter, request-weighted within each day and reported as the median of the last 7 daily values (text-generation requests only). Our comparison tool also calls models through OpenRouter, so this is closest to the speed you feel on this site. Models with no new value for over 14 days, or not on OpenRouter, fall back to Artificial Analysis's 72-hour median under a standard workload (about 10k input tokens, at least 1,500 output tokens) — the two are measured under different conditions, so compare them with care. Reasoning-mode measurements are not used.
Reading the score
Unit is tokens/second, higher is faster. A model that thinks (reasons) at length can have a high number here yet start answering late, so read it together with 'time to first token'. The same model varies by provider and time of day.

Try it yourself

Pick AIs and compare them.

Some AIs cannot be selected.

Rank

Top score210 tok/sMuse Spark 1.1
Models tested127
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Meta2026.07.09
    210 tok/s
  2. 2
    Arcee AI2026.04.01
    189 tok/s
  3. 3
    Mistral AI2026.04.28
    163 tok/s
  4. 4
    Google2025.09.25
    157 tok/s
  5. 5
    NVIDIA2025.12.14
    151 tok/s
  6. 6
    Google2026.03.03
    145 tok/s
  7. 7
    Google2026.07.21
    139 tok/s
  8. 8
    Meta2026.08.05
    139 tok/s
  9. 9
    Google2026.07.21
    136 tok/s
  10. 10
    Inference.net2026.04.16
    126 tok/s

9 of the 34 AIs released in the last 60 days have no score here yet · Perplexity Decider v1 27B, Clef Flash, Clef and 6 more

Source: OpenRouter · Artificial Analysis Data sources & removal requests

What to keep in mind

  • These are timing measurements. They vary with network, server load, and time of day, so your own experience can differ.
  • A hands-on comparison here is two AIs and one question, so it can differ from the official ranking. More rounds make it more accurate.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.