No. 02AI & LLM

LLM API Cost Calculator

Put in how many requests you make and how long they are, and compare the monthly bill across 10 current OpenAI, Anthropic and Google models — including what prompt caching saves.

lockRuns in your browser — nothing you type is uploaded.updateLast updated money_offFree, no sign-up
Your usage
%
Providers
Cheapest / month
$21.00
GPT-6 Luna · $0.00035 per request
Most expensive / month
$2,100
Claude Fable 5.1 · 60,000 requests
Monthly cost by model
  • GPT-6 Luna$21.00
  • Gemini 3.5 Flash-Lite$87.00
  • Gemini 3.8 Flash$157
  • Claude Haiku 4.5$210
  • GPT-6.1 Sol$420
  • Claude Sonnet 5.5$420
  • Gemini 3.1 Pro (preview)$468
  • Claude Opus 5.5$840
  • GPT-6 Astra$2,100
  • Claude Fable 5.1$2,100
Hover or tab to a model for its input/output split. * marks a note on that model.
Guide

How to use it

  1. 1
    Enter your request volume per day or per month. For a product that isn't live yet, estimate daily active users × requests per user.
  2. 2
    Enter average input and output tokens per request. Don't know them? Paste a typical prompt and answer into the token counter.
  3. 3
    If every request starts with the same long system prompt or document, set the share of input served from cache.
  4. 4
    Compare the bars, hover a model for its input/output split, and copy the link to share the scenario.
Worked examples

Examples

Support chatbot, 2,000 conversations a day

1,500 input and 400 output tokens each, about 60,000 requests a month. Gemini 3.5 Flash-Lite costs about $87 a month, Claude Haiku 4.5 $210, Claude Sonnet 5.5 and GPT-6.1 Sol $420 each, and Claude Opus 5.5 $840. The cheapest current model, GPT-6 Luna, comes to $21.

Same bot with an 80% cached system prompt

If 1,200 of the 1,500 input tokens are a fixed system prompt served from cache, Claude Sonnet 5.5 drops from $420 to about $290 a month and Claude Opus 5.5 from $840 to $566. The output side doesn't change, which is why caching helps least when answers are long.

Long-document analysis, 300 a day

250,000 input tokens and 1,000 output per request. Gemini 3.1 Pro crosses its 200K threshold and bills its higher long-prompt rate, about $9,160 a month — roughly the same as Claude Opus 5.5. Claude Haiku 4.5 can't take the request at all: its window is 200K tokens.

Method

How costs are calculated

For each model: input cost = input tokens × (uncached share × input price + cached share × cache-read price) and output cost = output tokens × output price, all per million tokens. Monthly cost = per-request cost × requests, with a month counted as 30 days.

Where a provider charges more for long prompts (Gemini 3.1 Pro above 200K input tokens), the higher rate is applied to the whole request once the threshold is crossed. Where no cache price is listed for a model here, cached tokens are billed at the full input rate.

ModelInput $/1MCached $/1MOutput $/1MContext
GPT-6 Astra$10$1$501050K
GPT-6.1 Sol$2$0.1$101050K
GPT-6 Luna$0.1$0.01$0.51050K
Claude Fable 5.1$10$0.25$501000K
Claude Opus 5.5$4$0.2$201000K
Claude Sonnet 5.5$2$0.2$101000K
Claude Haiku 4.5$1$0.1$5200K
Gemini 3.1 Pro (preview)$2 / $4—$12 / $181049K
Gemini 3.8 Flash$0.75—$3.751049K
Gemini 3.5 Flash-Lite$0.3—$2.51049K

Prices checked 2026-09-30 on each provider's pricing page. Two figures (x / y) mean standard / long-prompt rate.

Know before you rely on it

Limitations

  • info
    OpenAI bills prompts above 272K tokens at a higher long-context rate that isn't modelled here, so very long GPT-6 requests will cost more than shown.
  • info
    Gemini 3.8 Flash is on an introductory price until December 31, 2026; Google has announced it will double on January 1, 2027.
  • info
    Cache writes aren't counted. Anthropic charges a premium the first time a prompt is written to its cache, which matters when the cached part changes often.
  • info
    Reasoning models can spend extra “thinking” tokens that are billed as output. If you use extended reasoning, raise the output tokens to match what you see in real responses.

Sources: OpenAI API pricing · Anthropic pricing · Gemini API pricing

FAQ

Questions people ask

How is LLM API cost calculated?

add

Providers bill input tokens (everything you send: system prompt, context, conversation history) and output tokens (what the model writes) separately, each at a price per million tokens. Cost per request = input tokens × input price + output tokens × output price, divided by one million.

Why is output more expensive than input?

add

Generating a token takes much more compute than reading one, so every major provider charges more for output — typically four to six times the input price. For chatty products, output length often matters more than prompt length.

How much does prompt caching save?

add

Cached input is billed at a fraction of the normal input price — on current Claude models between 2.5% and 10% of it, on OpenAI's GPT-6 models 10%. If most of each request is a fixed system prompt or document, caching can cut the input part of the bill by more than half.

How many tokens is a word?

add

With current tokenizers, plain English prose runs about 1.1–1.3 tokens per word, or four to five characters per token. Code, non-English text and heavy formatting use more. Paste real prompts into the token counter to measure instead of guessing.

Are these the prices I'll actually pay?

add

They're the providers' standard pay-as-you-go API rates on the date shown. Batch processing, committed-use discounts, regional endpoints and cloud marketplaces (Bedrock, Vertex AI, Azure) can be cheaper or more expensive, and taxes aren't included.

From the directory

AI Coding Assistants worth a look

Launching an AI tool?

List it free on LaunchBoosts — it goes live the minute you submit and enters this week's Launch Race, where the top three win Product of the Week and a dofollow link.