Skip to content
Browse tools

Token counter & API cost estimator

Find out what your prompt actually costs.

Paste any prompt, document or code file. See how it breaks into tokens, how much of the context window it fills, and what it costs across 33 models from ten providers.

AI Token Calculator breaking a prompt into tokens and showing context usage and what the model call costs
33 models 10 providers runs in your browser nothing uploaded

Your text

0 characters
Try

Token breakdown

Tokens appear here as you type. Each shaded block is one token.
—

Step two

What it costs to run

Your input is only half the bill. Pick a scenario — one call, a whole conversation, or a month of traffic — and every number on the page follows.

Roughly 500 tokens ≈ a 375-word answer. A long report runs 2,000–4,000.


Cached reads bill at roughly 10% of the normal input rate on most providers.

Fixed indicative rate, not a live feed.

Running total

Input tokens per call0
Output tokens per call500
Input cost$0.00
Output cost$0.00
Cost per call$0.00
Total for 1 call $0.00
Per 1,000 calls
$0.00
Per 100,000 calls
$0.00
Per 1M calls
$0.00

Priced on — at — per million tokens.

Cheapest five for this workload

Side by side

Every model, same prompt

The same text priced across OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, xAI, Cohere, Amazon and Alibaba. Prices are per million tokens and every cell is editable — type your own negotiated rate and the table recalculates.

Model↕ Context↕ Est. input tokens Input $/M↕ Output $/M↕ Cost per call↕ Total Select model
—
Check before you budget. List prices move and providers add tiers for long context, batch mode and caching. Defaults here were reviewed in May 2026 — treat them as a starting point and confirm against your provider's current pricing page. Open-weight models are shown at typical hosted rates, which vary by host.

Beyond plain text

Images and embeddings

Vision inputs and vector indexing are billed in tokens too, and they are where budgets usually break. Two calculators for the parts most estimators skip.

Image token calculator

Image dimensions (pixels)
× images
0tokens total

—

Resizing to the smallest image that still answers the question is the cheapest optimisation in any vision workload — halving each side cuts tokens by roughly four.

Embedding cost calculator

$0.00per month

—

Quick reference

Tokens to words, at a glance

The conversion people look up most. Figures are for ordinary English prose at roughly 1.33 tokens per word.

Tokens Words Characters A4 pages Reading time Feels like

The same meaning in different languages

Tokenisers are trained mostly on English text, so other scripts fragment into more pieces. These figures show roughly how many tokens 100 English-equivalent words become.

LanguageTokens per 100 wordsvs EnglishWhy

Rules of thumb for English

~4characters per token
~0.75words per token
~1.3tokens per word
~750words per 1,000 tokens

Background

How tokens work

A language model never sees your letters. Before anything happens, your text is chopped into tokens — the units the model reads, generates and bills you for.

Not words

A token is a fragment

Common words are usually a single token. Rare or long ones split apart: tokenization becomes token + ization. Spaces normally travel with the word that follows them.

Both directions

You pay twice

Input tokens cover your prompt, system message, chat history and attached files. Output tokens cover the reply — and output usually costs three to five times more per token.

Language matters

English is the cheap case

The same meaning in Hindi, Thai or Arabic can take two to three times more tokens. Chinese and Japanese sit near one token per character.

Ways to cut your token bill

Prompt

Trim the boilerplate

Long politeness, repeated instructions and duplicated examples bill at full rate on every call. Say it once, clearly.

Data

Send less context

Retrieve the three relevant paragraphs instead of pasting the whole manual. Strip HTML, minify JSON, drop base64 blobs.

Reply

Cap the output

Ask for a fixed format and set a max token limit. "Answer in under 100 words" is the cheapest instruction you can write.

Caching

Reuse the stable part

If a long system prompt is identical across calls, prompt caching bills those reads at roughly a tenth of the usual rate.

Routing

Right-size the model

Classification and extraction rarely need a frontier model. Sending the easy 80% to a small model is usually the biggest single saving.

History

Summarise old turns

Replace turns one to ten with a 200-token summary. In long chats this is worth more than every other optimisation combined.

Questions

Good to know

How many words is 1,000 tokens?

About 750 words of ordinary English, or roughly 4,000 characters — around two and a half pages of double-spaced text. That works out to about 1.3 tokens per word. Technical writing, code, tables and non-English text all use more tokens for the same word count. The reference table above converts the common sizes.

How accurate is this token count?

It's a close estimate, not the provider's own tokeniser. The page models the same rules real byte-pair tokenisers follow — whole common words as single tokens, leading spaces attached to words, long words split into fragments, digits grouped, punctuation separate — then applies a per-family adjustment. On ordinary English prose it typically lands within a few percent. Code, unusual formatting and non-Latin scripts drift further. For billing-exact numbers, read the usage field the API returns with every response.

Is my text sent anywhere?

No. Everything runs locally in your browser. There's no request to a server, no analytics on the content, and nothing stored when you close the tab. You can disconnect from the network and the calculator still works.

Why does the same text cost different amounts on different models?

Two reasons. Each model family has its own vocabulary, so the same sentence splits into a slightly different number of tokens. And each has its own price per million tokens. A model with a cheaper rate but a coarser tokeniser wins twice over.

What exactly counts as input?

Everything you send: the system prompt, every earlier message in the conversation, tool and function definitions, tool results, and any documents or images you attach. In a long chat the history is resent on every turn — the Conversation tab above shows how quickly that compounds.

How are images counted as tokens?

By dimensions, using a per-provider formula. OpenAI charges 85 base tokens plus 170 per 512-pixel tile at high detail. Anthropic approximates width × height ÷ 750. Google charges 258 tokens per 768-pixel tile. The same 1024×1024 image costs about 765 tokens on OpenAI, 1,400 on Anthropic and 1,032 on Google — a near-two-fold spread for identical input. The image calculator above does the arithmetic for any size.

What is prompt caching and how much does it save?

Providers can store the unchanging prefix of your prompt — typically a long system message or a fixed document — and re-read it at a large discount, usually around 10% of the normal input rate. There's often a small premium to write the cache initially. It pays off when the same prefix is reused within minutes, which describes most production chat and retrieval systems.

What is a context window?

The maximum number of tokens the model can hold at once, counting your input and its reply together. Exceed it and the request fails or the oldest part of the conversation is dropped. The meter in the estimate panel shows how close you are for the model you picked.

Can I count tokens for a whole file?

Yes — use Upload file or drag a text file onto the input box. Plain text, Markdown, JSON, CSV and source code all work. Large files may take a moment to render the visual breakdown, so the display caps at the first 1,500 tokens with an option to expand.

Why is output priced higher than input?

Reading your prompt happens in one parallel pass. Generating a reply happens one token at a time, each step requiring a full forward pass through the model, so it uses far more compute per token. Providers price that difference in — which is why capping reply length saves more than trimming prompts.

Do reasoning models cost more than the price suggests?

Often, yes. Models that think before answering bill their internal reasoning as output tokens, even though you never see it. A 200-word visible answer can carry several thousand reasoning tokens behind it. If you're budgeting for a reasoning model, raise your output estimate substantially rather than trusting the visible reply length.

Where do these prices come from?

They are published list prices for each provider's API, reviewed in May 2026 and stated in US dollars per million tokens. Open-weight models such as Llama, Mistral Small, DeepSeek and Qwen have no single official price, so the figures shown are typical rates from hosted inference providers. Every price cell is editable — enter your own contracted rate and the whole page recalculates.

Estimates only. Confirm rates and counts with your provider before you commit to a budget. Prices reviewed May 2026. Runs entirely in your browser