What a token actually is
A model does not see words or characters. It sees tokens, produced by a tokeniser trained on real text. For English prose a useful average is:
1 token ≈ 0.75 words ≈ 4 characters
So a 1,000-word document is roughly 1,330 tokens, and "one million tokens" — the unit APIs are priced in — is about 750,000 words, roughly a 400-page book.
The average is useful for planning and dangerous as a guarantee, because the tokeniser splits on whatever units reduce the total. Three cases behave differently:
- Code and JSON cost about twice as much per character. Punctuation, indentation and symbols each become their own token. A JSON object of 18 characters can be ten tokens; a paragraph of the same length might be four.
- Non-English text costs 2-3x more. A tokeniser trained mostly on English encodes Chinese characters at roughly one token each, where English averages four characters per token. The same content can differ by 2-3x across languages, which matters a great deal if you are building a multilingual product.
- Numbers split for arithmetic reasons. The tokeniser learned to keep digits separable so models can do arithmetic, so a 15-digit ID may become five tokens. If your prompts carry timestamps, user IDs or hashes, you are paying for them twice — once in the tokens, once in the chance the model misreads them.
Why the count matters more than the price
Pricing is per token, but the constraint that actually breaks applications is the context window — the maximum number of tokens the model will read at once, counting both the input and the reply.
The arithmetic people get wrong is forgetting that the window is shared. A 128k model with a 100k-token document leaves 28k for the conversation, the system prompt and the answer. Four turns of a 5k-token conversation later, the request fails — not because anything was wrong with the document, but because the history consumed the budget.
This is why summarising or windowing history is not an optimisation but a requirement. The cost of conversation grows with the square of the turns if you resend everything, and a 20-turn chat on a 5k base is 100k tokens of history alone.
Input versus output, and the asymmetry that decides your bill
Output tokens are commonly priced 3-5x higher than input. The mechanism is simple: reading a token needs one forward pass, while generating each subsequent token needs the model invoked again.
Worked example at plausible rates — check current pricing, it moves quickly:
- 2,000 input tokens at $0.15 per million = $0.0003
- 800 output tokens at $0.60 per million = $0.00048
- Total: $0.00078 per call
Note that the output dominates despite being less than half the volume. Which side dominates depends entirely on the shape of your application: a chat with a long system prompt and short replies inverts the picture, and there the input is ten times the output.
Reducing the count, in order of effect
- Cut retrieved context. Most applications send far more than the question needs. Going from 50 KB to 2 KB of retrieved documents is usually the single largest saving available, and it improves quality by removing distraction.
- Cache a stable prefix. Most providers charge 10-25% of the normal input rate for a cached portion. A 1,500-token system prompt resent every call is 60% of a short request's cost, for free, if the prefix is cacheable. It needs a minimum length and entries expire, so a prefix that changes hourly never hits.
- Strip structure you do not need. Timestamps, IDs, hashes and base64 blobs in a prompt are pure cost. Move them to a tool call or a lookup.
- Cap the output. Asking for "a thorough answer" on a question needing a sentence multiplies the expensive side. Capping max tokens is the cheapest change available and usually improves quality.
- Cap the thinking budget on reasoning models. Extended reasoning produces hidden tokens that are billed as output. For straightforward tasks a low budget is both faster and much cheaper.
What this calculator does not do
It cannot count exactly, because each provider uses its own tokeniser and they are not interchangeable — the same text can differ by 10% or more between models, and more for code and non-English text. Only the provider's own endpoint gives an exact count.
What it does give is the right order of magnitude, the density diagnosis (is this text unusually token-dense, and why), and the arithmetic to compare options. That is enough to decide whether a design fits a context window and roughly what it will cost, which is what you need before writing the code. The AI cost calculator takes this further by comparing models and routing strategies.