Core Concepts
Tokens per Second in plain English.
Also known as: TPS,output speed,generation speed,throughput
The one-sentence version
The speed at which a model produces text, measured in tokens generated each second; the main number behind how fast an AI feels.
Tokens per second measures how quickly a model generates output. A token is roughly three-quarters of a word, so 50 tokens per second is about 40 words per second, comfortably faster than reading. Typical frontier models through their own APIs deliver 30 to 100 tokens per second; specialist inference hardware from Groq and Cerebras reaches several hundred to a few thousand on open-weight models. The number matters for different reasons in different products. In chat, anything above reading speed feels fine. In agents and reasoning models, which may generate tens of thousands of tokens of hidden thinking before answering, a 10x speed difference is the gap between a two-minute wait and twelve seconds. In voice, it must be fast enough to stay ahead of speech. Tokens per second is not the whole latency story: time to first token, which covers processing your prompt, can dominate for long inputs. Providers publish both, and benchmark sites measure them independently because vendor figures usually reflect ideal conditions.