AI

AI API Cost Calculator

Per-token list prices are the least useful number in AI planning. This models what a workload actually costs per month, then shows what caching, routing and batching do to it — you enter the rates, so nothing here goes stale.

Result
—
Monthly cost
—
Input vs output split
—
Per call and per million calls
—
Monthly token volume
—
With prefix caching
—
With tighter retrieval
—
With model routing
—
With batching
—
Best single lever
—
Where to look first
—

List price is the least useful number

Per-million-token pricing is what vendors publish, and it is almost never what you pay. Three things sit between the two:

  • Calls per request. A typical application is not one call. Classify, retrieve, generate, validate — and retry on validation failure sometimes. Teams that model only the headline call are typically off by 3-5x.
  • Tokens per call. The visible question is rarely most of the input. System prompts, tool definitions and retrieved documents are all billed, and the user's question is often the smallest part.
  • Which side dominates. Output is priced 3-5x input, so 500 output tokens can cost more than 10,000 input tokens. Which side dominates decides which optimisation matters.

The four levers, in order of typical effect

1. Trim the input. This is almost always the biggest single win, and it is a quality improvement rather than a compromise. Most applications send far more context than the question needs; cutting retrieved documents from 50 KB to 2 KB typically removes more cost than any other change, and it also reduces the chance the model answers from an irrelevant passage.

2. Cache a stable prefix. Providers that support prompt caching charge 10-25% of the normal input rate for the cached portion. This pays off when a large prefix is genuinely reused — a long system prompt, a few-shot block, a document that does not change. It requires a minimum prefix length and entries expire after minutes to an hour, so a prefix that changes every request never hits the cache. Measure your hit rate before assuming a saving.

3. Route by difficulty. Most production traffic is easy: classification, extraction, formatting, simple questions. Send those to a small model and reserve the frontier model for the hard remainder. A tiered setup where a small model handles 70-80% of traffic typically cuts total spend by 60-80% at equal or better quality, because the quality that matters lives in the hard 20%.

4. Batch. Providers offer 30-50% discounts for requests submitted together with a longer turnaround. For overnight processing, evaluation runs, or any pipeline with a queue, this is close to free money. It is not appropriate for interactive work, where the latency is the product.

A realistic worked model

A document assistant: 20,000 requests a month, 10,000 input tokens each (4,000 of user question plus 6,000 of retrieved context), 500 output tokens, at $0.15/M in and $0.60/M out.

  • Input: 20,000 × 10,000 ÷ 1e6 × 0.15 = $30.00
  • Output: 20,000 × 500 ÷ 1e6 × 0.60 = $6.00
  • Total: $36.00 per month

Two observations that generalise. The retrieved context is 60% of the input volume, which makes retrieval precision the biggest lever rather than model choice. And $36 is a small number — which means the engineering work to halve it is rarely worth it, until volume grows. At 200,000 requests the same arithmetic gives $360, and at 2 million it gives $3,600, where the optimisation work pays for itself many times over. Model cost at the volume you have, not the volume you hope for.

The comparison that is not about price

Price per token is the least interesting difference between models. Two others dominate in practice:

Cost per successful task. A small model that needs three retries, emits malformed JSON requiring a repair pass, or fails a validation check costs more than a large model that succeeds first time. The comparison that matters is total system cost per completed task, which includes parsing, retries and human review of failures.

Latency and its knock-on effects. A slower model is not just less pleasant; in an agentic loop where one call follows another, latency multiplies across steps. For interactive products the slow path is often the wrong path regardless of its unit price.

What this calculator does not do

It does not include embeddings, reranking, image or audio input, tool-definition tokens, or rate-limit-induced queueing — all of which appear on real invoices and are easy to forget when modelling. It also does not know your prices; you enter them, because they change frequently and a stale hard-coded table is worse than none.

Frequently asked questions

1. How much does an LLM API cost per month?

Multiply requests by (input tokens + output tokens), apply the per-million rates, then multiply by the number of calls per request and the retry rate. A 20,000-request application at 10,000 input and 500 output tokens with $0.15/$0.60 rates costs about $36 per month.

2. What is the cheapest way to reduce AI API cost?

Trimming the input is almost always the biggest single win — most applications send far more context than the question needs, and cutting it is a quality improvement rather than a compromise. Prefix caching and model routing come next, and batching is close to free for offline work.

3. Does prompt caching really save money?

When a large prefix is stable across requests, yes — typically 10-25% of the input rate for the cached portion. It requires a minimum prefix length and entries expire, so short or changing prompts never hit. Measure the cache hit rate before assuming a saving.

4. Is a cheaper model always cheaper overall?

No. A smaller model that needs retries, produces malformed output requiring a parser, or fails validation can cost more than a larger model that succeeds first time. The right comparison is cost per successful task, not cost per token.

Related tools