Token counter & API cost estimator
Find out what your prompt actually costs.
Paste any prompt, document or code file. See how it breaks into tokens, how much of the context window it fills, and what it costs across 33 models from ten providers.
Your text
Token breakdown
Step two
What it costs to run
Your input is only half the bill. Pick a scenario — one call, a whole conversation, or a month of traffic — and every number on the page follows.
Roughly 500 tokens ≈ a 375-word answer. A long report runs 2,000–4,000.
Chat history is resent on every turn, so cost grows quadratically. Your pasted text is treated as the system prompt.
Input per request comes from the text you pasted above.
Cached reads bill at roughly 10% of the normal input rate on most providers.
Fixed indicative rate, not a live feed.
Running total
- Per 1,000 calls
- $0.00
- Per 100,000 calls
- $0.00
- Per 1M calls
- $0.00
Priced on — at — per million tokens.
Cheapest five for this workload
Side by side
Every model, same prompt
The same text priced across OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, xAI, Cohere, Amazon and Alibaba. Prices are per million tokens and every cell is editable — type your own negotiated rate and the table recalculates.
| Model↕ | Context↕ | Est. input tokens | Input $/M↕ | Output $/M↕ | Cost per call↕ | Total | Select model |
|---|
Beyond plain text
Images and embeddings
Vision inputs and vector indexing are billed in tokens too, and they are where budgets usually break. Two calculators for the parts most estimators skip.
Image token calculator
—
Embedding cost calculator
—
Quick reference
Tokens to words, at a glance
The conversion people look up most. Figures are for ordinary English prose at roughly 1.33 tokens per word.
| Tokens | Words | Characters | A4 pages | Reading time | Feels like |
|---|
The same meaning in different languages
Tokenisers are trained mostly on English text, so other scripts fragment into more pieces. These figures show roughly how many tokens 100 English-equivalent words become.
| Language | Tokens per 100 words | vs English | Why |
|---|
Rules of thumb for English
Background
How tokens work
A language model never sees your letters. Before anything happens, your text is chopped into tokens — the units the model reads, generates and bills you for.
Not words
A token is a fragment
Common words are usually a single token. Rare or long ones split apart: tokenization becomes token + ization. Spaces normally travel with the word that follows them.
Both directions
You pay twice
Input tokens cover your prompt, system message, chat history and attached files. Output tokens cover the reply — and output usually costs three to five times more per token.
Language matters
English is the cheap case
The same meaning in Hindi, Thai or Arabic can take two to three times more tokens. Chinese and Japanese sit near one token per character.
Ways to cut your token bill
Prompt
Trim the boilerplate
Long politeness, repeated instructions and duplicated examples bill at full rate on every call. Say it once, clearly.
Data
Send less context
Retrieve the three relevant paragraphs instead of pasting the whole manual. Strip HTML, minify JSON, drop base64 blobs.
Reply
Cap the output
Ask for a fixed format and set a max token limit. "Answer in under 100 words" is the cheapest instruction you can write.
Caching
Reuse the stable part
If a long system prompt is identical across calls, prompt caching bills those reads at roughly a tenth of the usual rate.
Routing
Right-size the model
Classification and extraction rarely need a frontier model. Sending the easy 80% to a small model is usually the biggest single saving.
History
Summarise old turns
Replace turns one to ten with a 200-token summary. In long chats this is worth more than every other optimisation combined.
Questions
Good to know
How many words is 1,000 tokens?
About 750 words of ordinary English, or roughly 4,000 characters — around two and a half pages of double-spaced text. That works out to about 1.3 tokens per word. Technical writing, code, tables and non-English text all use more tokens for the same word count. The reference table above converts the common sizes.
How accurate is this token count?
It's a close estimate, not the provider's own tokeniser. The page models the same rules real byte-pair tokenisers follow — whole common words as single tokens, leading spaces attached to words, long words split into fragments, digits grouped, punctuation separate — then applies a per-family adjustment. On ordinary English prose it typically lands within a few percent. Code, unusual formatting and non-Latin scripts drift further. For billing-exact numbers, read the usage field the API returns with every response.
Is my text sent anywhere?
No. Everything runs locally in your browser. There's no request to a server, no analytics on the content, and nothing stored when you close the tab. You can disconnect from the network and the calculator still works.
Why does the same text cost different amounts on different models?
Two reasons. Each model family has its own vocabulary, so the same sentence splits into a slightly different number of tokens. And each has its own price per million tokens. A model with a cheaper rate but a coarser tokeniser wins twice over.
What exactly counts as input?
Everything you send: the system prompt, every earlier message in the conversation, tool and function definitions, tool results, and any documents or images you attach. In a long chat the history is resent on every turn — the Conversation tab above shows how quickly that compounds.
How are images counted as tokens?
By dimensions, using a per-provider formula. OpenAI charges 85 base tokens plus 170 per 512-pixel tile at high detail. Anthropic approximates width × height ÷ 750. Google charges 258 tokens per 768-pixel tile. The same 1024×1024 image costs about 765 tokens on OpenAI, 1,400 on Anthropic and 1,032 on Google — a near-two-fold spread for identical input. The image calculator above does the arithmetic for any size.
What is prompt caching and how much does it save?
Providers can store the unchanging prefix of your prompt — typically a long system message or a fixed document — and re-read it at a large discount, usually around 10% of the normal input rate. There's often a small premium to write the cache initially. It pays off when the same prefix is reused within minutes, which describes most production chat and retrieval systems.
What is a context window?
The maximum number of tokens the model can hold at once, counting your input and its reply together. Exceed it and the request fails or the oldest part of the conversation is dropped. The meter in the estimate panel shows how close you are for the model you picked.
Can I count tokens for a whole file?
Yes — use Upload file or drag a text file onto the input box. Plain text, Markdown, JSON, CSV and source code all work. Large files may take a moment to render the visual breakdown, so the display caps at the first 1,500 tokens with an option to expand.
Why is output priced higher than input?
Reading your prompt happens in one parallel pass. Generating a reply happens one token at a time, each step requiring a full forward pass through the model, so it uses far more compute per token. Providers price that difference in — which is why capping reply length saves more than trimming prompts.
Do reasoning models cost more than the price suggests?
Often, yes. Models that think before answering bill their internal reasoning as output tokens, even though you never see it. A 200-word visible answer can carry several thousand reasoning tokens behind it. If you're budgeting for a reasoning model, raise your output estimate substantially rather than trusting the visible reply length.
Where do these prices come from?
They are published list prices for each provider's API, reviewed in May 2026 and stated in US dollars per million tokens. Open-weight models such as Llama, Mistral Small, DeepSeek and Qwen have no single official price, so the figures shown are typical rates from hosted inference providers. Every price cell is editable — enter your own contracted rate and the whole page recalculates.
What else can you check?
Everything here runs in your browser the same way this calculator does — no upload, no account, no waiting.