The number that matters is cost per user
Per-call pricing flatters chat products. What determines whether a chat feature is viable is cost per active user per month, and that number is dominated by something most estimates omit: the conversation is resent every turn.
For a 25-turn conversation, the total input is not 25 × (system + question). It is the system prompt 25 times, plus history that grows by one turn each time. The average history per call is (M+1)/2 × per-turn tokens, so:
- 5 turns: average history 3 × per-turn
- 25 turns: average history 13 × per-turn — 4.3x more per call than a 5-turn chat
- 50 turns: average history 25.5 × per-turn
Why the cost is superlinear in message count
This is the single most useful thing to understand about chat economics. Total input tokens for an M-turn conversation is M × S + H × M(M+1)/2 — the history term is quadratic in M.
Worked comparison, 800-token system prompt, 1,200 tokens of history per turn, 180 output:
- 5 turns/user: 5 × 800 + 1,200 × 15 = 22,000 input tokens
- 25 turns/user: 25 × 800 + 1,200 × 325 = 410,000 input tokens — 18.6x for 5x the turns
- 50 turns/user: 50 × 800 + 1,200 × 1,275 = 1,570,000 input tokens — 71x
The practical consequence: shortening conversations is worth more than shortening prompts. A product that caps a thread at 10 turns and summarises the rest pays a fraction of one that lets threads run to 50.
The three costs nobody models
1. Retries. A retry rate of 8% looks small and costs about 7.4% of the bill — pure waste. Common causes are strict structured-output requirements, over-long outputs hitting the cap, and transient rate limits. Each is separately fixable: relax the schema, cap output explicitly, add backoff.
2. Tool round-trips. A single user-visible answer may involve several model calls: one to decide which tool to use, one to run it, one to interpret the result. A 25-turn conversation with tools can be 3-4x the calls the naive count suggests. Tool definitions are also billed as input on every single call, so an agent with 20 tools carries 2,000+ tokens of schema per request before the user types anything.
3. Abandoned threads. Users start conversations they do not finish. If the average thread runs 40% of its turns before being abandoned, a per-user estimate based on completed conversations overstates nothing — but it does mean the realised figure is lower than modelled, which is the good kind of surprise.
Where the money actually is
For most chat products, in descending order:
- Resent history. Summarise or window. This is the biggest lever and it also improves quality, since a long history degrades relevance.
- Tool-definition tokens. Ship fewer tools, or make them retrievable so only the relevant schemas are sent.
- Output length. Cap max_tokens and ask for concise answers. Output is the expensive side at 3-5x the input rate.
- Retries. Instrument them; each category has a different fix.
- Model tier. Route simple turns to a small model. In a chat product most turns are simple.
What this calculator does not do
It does not model tool calls, embeddings, image input, or caching — all of which appear on real invoices. It also does not include subscription pricing: a flat monthly fee per user changes the economics completely, and the comparison worth making is API cost per user against the subscription price that would cover it. For a product with heavy usage, the API bill is why flat-rate subscriptions exist.