AI API Costs Explained: Tokens, Caching and the Bill That Surprises

October 4, 2026 · 10 min read

AI API pricing looks simple — a number per million tokens — and the bills can look inexplicable. Almost every surprise traces back to three things: what a token actually is, the asymmetry between reading and writing, and the fact that a single user request becomes several model calls.

A token is about three quarters of a word

Models do not see words or characters; they see tokens, produced by a tokeniser trained on real text. For English prose, a useful rule is:

1 token ≈ 0.75 words ≈ 4 characters

So a 1,000-word document is roughly 1,330 tokens, and pricing quoted "per million tokens" means a million tokens is about 750,000 words — roughly a 400-page book. The token calculator estimates this for a given input.

Three things make real text deviate:

  • Code costs more per character. Punctuation, indentation and symbols each become tokens, and code is dense in them. A function can be 1.5-2x the token density of prose.
  • Non-English text costs more. A tokeniser trained mostly on English encodes Chinese characters roughly one token each, whereas English averages 4 characters per token. The same content can differ by 2-3x across languages.
  • Numbers split unpredictably. A long ID or a timestamp may tokenize into several pieces. This is a real cost issue in production, where IDs and timestamps are everywhere in structured data.

Output costs multiples of input, and why

This is the asymmetry that surprises most people. Processing a prompt is one forward pass; generating a token requires the model to be invoked again for each subsequent token. Providers price that difference in, and output is commonly 3-5x the input rate.

Worked example at plausible 2026 list prices — always check current pricing, it moves fast:

  • Input: 2,000 tokens at $0.15 per million = $0.0003
  • Output: 800 tokens at $0.60 per million = $0.00048
  • Total: $0.00078 per call — well under a tenth of a cent

The output dominates despite being less than half the volume. A chat application with long system prompts and short replies inverts the picture: the input is 8,000 tokens costing $0.0012 and the output 200 tokens costing $0.00012, so now the input dominates by 10x. Which side dominates depends entirely on the shape of your application, and the optimisation strategy depends on it too.

A user request is rarely one model call

The per-call cost is rarely what the customer pays for. A typical application does something like:

  1. Classify the request (small, cheap model) — 0.0001
  2. Retrieve relevant context (embedding model, usually per token, very cheap)
  3. Generate the answer (large model) — 0.0008
  4. Validate or self-check the output (another large-model call) — 0.0004
  5. Occasionally retry on validation failure — probabilistic

Multiply that by the request rate and the retry rate, and the per-request figure of 0.0001 dollars becomes 0.002 or more. When modelling cost, multiply by the number of calls per user request and by the expected number of attempts, then multiply by requests. Teams that cost-model only the headline call are typically off by 3-5x.

Three ways to cut the bill without cutting quality

1. Prompt caching. Most providers will cache a stable prefix and charge 10-25% of the normal input rate for the cached portion. This is the single biggest lever when you have a long system prompt, a few-shot block, or a large document that does not change between calls. It only pays when the prefix is genuinely stable — a cache that misses every time costs nothing and helps nothing.

2. Model routing. Use a large model only for the requests that need it. Most production traffic is easy: classification, extraction, simple questions. A tiered setup where a small model handles 70-80% of traffic and a large model handles the rest typically cuts total spend by 60-80% at equal or better quality, because the quality that matters lives in the hard 20%.

3. Batching. Providers give a discount — often 50% — for requests submitted together as a batch with a longer turnaround. For offline work, overnight processing, or any pipeline with a queue, this is close to free money. The AI cost calculator models these options side by side.

What the pricing pages do not tell you

  • Tokens are not only the visible text. Tool definitions, function schemas, retrieved documents and conversation history are all billed as input. A 20-turn conversation carries the entire history every time, so its cost grows with the square of the turns unless you summarise.
  • Repeated context is the most common waste. A 50 KB document in every request costs the same whether it is needed for the answer or not. Retrieval that selects 2 KB instead of 50 KB is the largest single saving available to most applications.
  • Output length is a design choice. Asking for "a short answer" or capping max tokens is often the cheapest change available, and it also tends to improve quality by removing the incentive to pad.
  • Free tiers have rate limits, not free usage. A free tier is for testing. Building a product on one produces a bill the first time it succeeds.

A realistic monthly model

Abstract per-call pricing does not survive contact with an application. A worked example, using plausible 2026 list prices:

  • Application: a document assistant. 2,000 requests per month, each with a 4,000-token user question plus 6,000 tokens of retrieved context, a 500-token answer, one validation call, and a 15% retry rate.
  • Per request: main call input 10,000 tokens at $0.15/M = $0.0015, output 500 at $0.60/M = $0.0003 → $0.0018. Validation call at 20% of that = $0.0004. With a 15% retry rate the expected cost is $0.0025.
  • Monthly: 2,000 × $0.0025 = $5.

Two observations. The retrieved context is 60% of the input volume — which makes retrieval quality the biggest cost lever, not the model choice. And 2,000 requests is a small product; at 200,000 the same arithmetic gives $500, and the fixed engineering work (prompt optimisation, caching, routing) is what turns a $5 line item into a meaningful saving.

What a token actually is, more precisely

The "0.75 words" rule is a useful average and a poor guarantee. A tokeniser trained on web text splits on whatever units reduce total tokens, which produces results that feel arbitrary in the cases that matter:

  • JSON and code split aggressively. {"user_id": 12345} is roughly 10 tokens for 18 characters, about 2x the density of English prose. A tool-calling application whose entire output is JSON is paying a premium on the response side.
  • Whitespace runs merge differently by language. English spaces often attach to the following word; languages without spaces tokenize much less efficiently, and the cost of the same content can differ by 2-3x.
  • Numbers are split for arithmetic reasons — the tokeniser learned to keep digits separable so models can do arithmetic. A 15-digit ID may become 5 tokens, while the same 15 characters as letters might be 3.

This is why the same application costs noticeably less after you strip timestamps, IDs and hashes from the prompt. Every one of them is being paid for twice: once in the input, and again in the model's chance to misread it.

Where the waste actually is

In production systems, in descending order of typical impact:

  1. Over-broad retrieval. Fetching 50 KB of documents when 2 KB would answer the question. Retrieval precision is a cost lever before it is a quality lever.
  2. Untruncated conversation history. A 20-turn conversation sends all prior turns every time, so total input grows with the square of the turns. Summarising or windowing is the fix, and it also improves quality by removing distraction.
  3. No caching of a stable prefix. A 1,500-token system prompt and few-shot block resent on every request is 60% of a short request's cost, for free, if the prefix is cacheable.
  4. Oversized generation. Asking for "a thorough answer" on a question needing a sentence multiplies the output cost, which is the expensive side. Capping max_tokens is the cheapest change available and usually improves quality too.
  5. Model choice per difficulty. Using a frontier model for classification and extraction. Routing is where the largest multiples live — 60-80% of traffic typically needs a small model.

What the pricing pages do not say

  • Tool definitions and schemas are billed input. A function-calling app with 20 tools sends 2,000 tokens of tool definitions on every single call, before the user says anything.
  • Embeddings and reranking are separate bills with their own per-token pricing, usually far cheaper than generation but not zero, and they are easy to forget when modelling.
  • Caching has a TTL and a minimum length. Most providers require a prefix of at least a few hundred tokens, and cached entries expire after 5-60 minutes. A prefix that changes every hour never gets a cache hit.
  • Rate limits are not pricing but they shape design. A lower tier can mean a lower rate limit, and hitting it means queueing — which trades latency for cost invisibly, in the form of a slower product.

Frequently asked questions

How do I estimate my monthly LLM API cost?

Multiply requests per month by (input tokens + output tokens) per request, apply the per-million rates, then multiply by the number of model calls per request and the expected retry rate. Remember that retrieved context and tool definitions count as input even though the user did not type them.

Why is my token count higher than I expected?

Because a token averages 4 characters of English prose but much less for code, JSON, numbers and non-English text. Structural characters each become their own token, and digits are split so the model can do arithmetic on them.

Does prompt caching really help?

When a large prefix is stable across requests, yes — typically 10-25% of the input rate for the cached portion. It requires a minimum prefix length and entries expire, so short or changing prompts never hit. Measure the cache hit rate before assuming a saving.

What is the single biggest cost optimisation?

Retrieval precision, usually. Most applications send far more context than the question needs. Cutting retrieved context from 50 KB to 2 KB typically removes more cost than switching model tier, and it also improves answer quality by removing distraction.

Frequently asked questions

What is a token?

The unit models actually bill in, roughly 0.75 English words or about 4 characters. A 1000-word document is around 1,300 tokens. The tokeniser also splits on structure — punctuation, numbers and code fragments often become separate tokens, which is why code is more expensive per character than prose.

Why do output tokens cost more than input tokens?

Because generating a token requires running the model, while reading one only requires the computation up to that point. On a per-token basis output is commonly 3-5x the input price, so a chat application that returns long answers costs far more than one with long inputs and short replies.

How much does prompt caching save?

Providers that support it typically charge 10-25% of the normal input rate for the cached portion. This only pays off when a large, stable prefix is reused across calls — a long system prompt or a document that does not change. If the prefix changes every request, caching saves nothing.

Is a cheaper model always cheaper in total?

No. A smaller model that needs three retries, produces malformed output requiring a parser, or gets rejected by a validator can cost more than a larger model that succeeds first time. The right comparison is cost per successful task, not cost per call.

Related guides