How to use it
- 1Paste or type your text. The count updates as you type; the tokenizer loads in the background the first time.
- 2Set how many tokens you expect back in the answer — it counts toward both the context window and the cost.
- 3Check the table: the bar shows how much of each model's context the call uses, and the cost columns show one call and 1,000 calls.
- 4Open See the tokens to see exactly where the text is split — useful for spotting wasteful formatting.
Examples
A short support prompt
The built-in sample — a two-line system instruction plus a user question — is 53 tokens for 248 characters — about 4.7 characters per token. With a 500-token answer, one call on Claude Sonnet 5.5 costs about half a cent.
Formatting costs tokens
The same product data as pretty-printed JSON uses far more tokens than as a compact list, because every indent, quote and brace is counted. Paste both versions and compare — trimming formatting in a prompt you send thousands of times a day is one of the cheapest optimisations there is.
Will this document fit?
A 300-page book comes to roughly 100,000–130,000 tokens at 1.1 tokens per word. On Claude Haiku 4.5's 200K window that's up to two-thirds full before your instructions and the answer; paste a second book and its bar turns red, while the 1M-token models still show plenty of headroom.
How the count and costs are produced
Text is encoded with the byte-pair encoding o200k_base using the open-source gpt-tokenizer library, running in a Web Worker so long pastes don't freeze the page. Characters are counted as Unicode code points; words are runs of non-whitespace.
Context use = (text tokens + expected output tokens) ÷ the model's context window. Cost per call = text tokens × input price + output tokens × output price. Prices were checked on 2026-09-30; the LLM API cost calculator lists them all and models caching and monthly volume.
Limitations
- infoRows marked ≈ reuse the o200k count for models with different tokenizers. Treat them as estimates, and verify with the provider's own token-counting API before relying on a tight fit.
- infoImages, audio, PDFs and tool definitions are billed as tokens too but can't be counted from pasted text.
- infoThe token view shows the first 400 tokens. Pieces that look empty are spaces or line breaks.
Sources: OpenAI pricing · Anthropic pricing · Gemini pricing
Questions people ask
What is a token?
add
A token is the unit a language model reads and writes: often a whole short word, sometimes part of a longer word, a punctuation mark or a space. “Launching” might be one token, while an unusual product name can be split into three or four. Models have context limits and prices counted in tokens, not words.
Which tokenizer does this counter use?
add
OpenAI's o200k_base encoding, the one used by GPT-4o, the o-series and GPT-5 models, via the open-source gpt-tokenizer library. For those models the count is exact. Newer or other models may split text slightly differently.
Are the counts accurate for Claude and Gemini?
add
They're close estimates, not exact. Anthropic and Google use their own tokenizers, and Anthropic notes that its newer models produce roughly 30% more tokens than its older ones for the same text. For an exact figure, use Anthropic's count_tokens endpoint or Gemini's countTokens method.
How many tokens is 1,000 words?
add
For plain English prose, about 1,100 tokens with o200k_base — we measured 1.1 tokens per word on a sample essay. The often-quoted “1.3 tokens per word” comes from older tokenizers. Code, JSON, tables and non-Latin scripts use noticeably more. Paste your own text: the tokens-per-word figure shows how it compares.
Is my text sent anywhere?
add
No. The tokenizer is downloaded to your browser and runs in a background thread on your device. Nothing you paste is uploaded or stored.
Why does my prompt use more tokens in the API than here?
add
Chat APIs wrap each message with role markers and formatting tokens, and tool definitions, images and system prompts are counted too. Expect a few extra tokens per message on top of the raw text.
