List price is the least useful number
Per-million-token pricing is what vendors publish, and it is almost never what you pay. Three things sit between the two:
- Calls per request. A typical application is not one call. Classify, retrieve, generate, validate — and retry on validation failure sometimes. Teams that model only the headline call are typically off by 3-5x.
- Tokens per call. The visible question is rarely most of the input. System prompts, tool definitions and retrieved documents are all billed, and the user's question is often the smallest part.
- Which side dominates. Output is priced 3-5x input, so 500 output tokens can cost more than 10,000 input tokens. Which side dominates decides which optimisation matters.
The four levers, in order of typical effect
1. Trim the input. This is almost always the biggest single win, and it is a quality improvement rather than a compromise. Most applications send far more context than the question needs; cutting retrieved documents from 50 KB to 2 KB typically removes more cost than any other change, and it also reduces the chance the model answers from an irrelevant passage.
2. Cache a stable prefix. Providers that support prompt caching charge 10-25% of the normal input rate for the cached portion. This pays off when a large prefix is genuinely reused — a long system prompt, a few-shot block, a document that does not change. It requires a minimum prefix length and entries expire after minutes to an hour, so a prefix that changes every request never hits the cache. Measure your hit rate before assuming a saving.
3. Route by difficulty. Most production traffic is easy: classification, extraction, formatting, simple questions. Send those to a small model and reserve the frontier model for the hard remainder. A tiered setup where a small model handles 70-80% of traffic typically cuts total spend by 60-80% at equal or better quality, because the quality that matters lives in the hard 20%.
4. Batch. Providers offer 30-50% discounts for requests submitted together with a longer turnaround. For overnight processing, evaluation runs, or any pipeline with a queue, this is close to free money. It is not appropriate for interactive work, where the latency is the product.
A realistic worked model
A document assistant: 20,000 requests a month, 10,000 input tokens each (4,000 of user question plus 6,000 of retrieved context), 500 output tokens, at $0.15/M in and $0.60/M out.
- Input: 20,000 × 10,000 ÷ 1e6 × 0.15 = $30.00
- Output: 20,000 × 500 ÷ 1e6 × 0.60 = $6.00
- Total: $36.00 per month
Two observations that generalise. The retrieved context is 60% of the input volume, which makes retrieval precision the biggest lever rather than model choice. And $36 is a small number — which means the engineering work to halve it is rarely worth it, until volume grows. At 200,000 requests the same arithmetic gives $360, and at 2 million it gives $3,600, where the optimisation work pays for itself many times over. Model cost at the volume you have, not the volume you hope for.
The comparison that is not about price
Price per token is the least interesting difference between models. Two others dominate in practice:
Cost per successful task. A small model that needs three retries, emits malformed JSON requiring a repair pass, or fails a validation check costs more than a large model that succeeds first time. The comparison that matters is total system cost per completed task, which includes parsing, retries and human review of failures.
Latency and its knock-on effects. A slower model is not just less pleasant; in an agentic loop where one call follows another, latency multiplies across steps. For interactive products the slow path is often the wrong path regardless of its unit price.
What this calculator does not do
It does not include embeddings, reranking, image or audio input, tool-definition tokens, or rate-limit-induced queueing — all of which appear on real invoices and are easy to forget when modelling. It also does not know your prices; you enter them, because they change frequently and a stale hard-coded table is worse than none.