How to use it
- 1Enter your request volume per day or per month. For a product that isn't live yet, estimate daily active users × requests per user.
- 2Enter average input and output tokens per request. Don't know them? Paste a typical prompt and answer into the token counter.
- 3If every request starts with the same long system prompt or document, set the share of input served from cache.
- 4Compare the bars, hover a model for its input/output split, and copy the link to share the scenario.
Examples
Support chatbot, 2,000 conversations a day
1,500 input and 400 output tokens each, about 60,000 requests a month. Gemini 3.5 Flash-Lite costs about $87 a month, Claude Haiku 4.5 $210, Claude Sonnet 5.5 and GPT-6.1 Sol $420 each, and Claude Opus 5.5 $840. The cheapest current model, GPT-6 Luna, comes to $21.
Same bot with an 80% cached system prompt
If 1,200 of the 1,500 input tokens are a fixed system prompt served from cache, Claude Sonnet 5.5 drops from $420 to about $290 a month and Claude Opus 5.5 from $840 to $566. The output side doesn't change, which is why caching helps least when answers are long.
Long-document analysis, 300 a day
250,000 input tokens and 1,000 output per request. Gemini 3.1 Pro crosses its 200K threshold and bills its higher long-prompt rate, about $9,160 a month — roughly the same as Claude Opus 5.5. Claude Haiku 4.5 can't take the request at all: its window is 200K tokens.
How costs are calculated
For each model: input cost = input tokens × (uncached share × input price + cached share × cache-read price) and output cost = output tokens × output price, all per million tokens. Monthly cost = per-request cost × requests, with a month counted as 30 days.
Where a provider charges more for long prompts (Gemini 3.1 Pro above 200K input tokens), the higher rate is applied to the whole request once the threshold is crossed. Where no cache price is listed for a model here, cached tokens are billed at the full input rate.
| Model | Input $/1M | Cached $/1M | Output $/1M | Context |
|---|---|---|---|---|
| GPT-6 Astra | $10 | $1 | $50 | 1050K |
| GPT-6.1 Sol | $2 | $0.1 | $10 | 1050K |
| GPT-6 Luna | $0.1 | $0.01 | $0.5 | 1050K |
| Claude Fable 5.1 | $10 | $0.25 | $50 | 1000K |
| Claude Opus 5.5 | $4 | $0.2 | $20 | 1000K |
| Claude Sonnet 5.5 | $2 | $0.2 | $10 | 1000K |
| Claude Haiku 4.5 | $1 | $0.1 | $5 | 200K |
| Gemini 3.1 Pro (preview) | $2 / $4 | — | $12 / $18 | 1049K |
| Gemini 3.8 Flash | $0.75 | — | $3.75 | 1049K |
| Gemini 3.5 Flash-Lite | $0.3 | — | $2.5 | 1049K |
Prices checked 2026-09-30 on each provider's pricing page. Two figures (x / y) mean standard / long-prompt rate.
Limitations
- infoOpenAI bills prompts above 272K tokens at a higher long-context rate that isn't modelled here, so very long GPT-6 requests will cost more than shown.
- infoGemini 3.8 Flash is on an introductory price until December 31, 2026; Google has announced it will double on January 1, 2027.
- infoCache writes aren't counted. Anthropic charges a premium the first time a prompt is written to its cache, which matters when the cached part changes often.
- infoReasoning models can spend extra “thinking” tokens that are billed as output. If you use extended reasoning, raise the output tokens to match what you see in real responses.
Sources: OpenAI API pricing · Anthropic pricing · Gemini API pricing
Questions people ask
How is LLM API cost calculated?
add
Providers bill input tokens (everything you send: system prompt, context, conversation history) and output tokens (what the model writes) separately, each at a price per million tokens. Cost per request = input tokens × input price + output tokens × output price, divided by one million.
Why is output more expensive than input?
add
Generating a token takes much more compute than reading one, so every major provider charges more for output — typically four to six times the input price. For chatty products, output length often matters more than prompt length.
How much does prompt caching save?
add
Cached input is billed at a fraction of the normal input price — on current Claude models between 2.5% and 10% of it, on OpenAI's GPT-6 models 10%. If most of each request is a fixed system prompt or document, caching can cut the input part of the bill by more than half.
How many tokens is a word?
add
With current tokenizers, plain English prose runs about 1.1–1.3 tokens per word, or four to five characters per token. Code, non-English text and heavy formatting use more. Paste real prompts into the token counter to measure instead of guessing.
Are these the prices I'll actually pay?
add
They're the providers' standard pay-as-you-go API rates on the date shown. Batch processing, committed-use discounts, regional endpoints and cloud marketplaces (Bedrock, Vertex AI, Azure) can be cheaper or more expensive, and taxes aren't included.
