What is a token, and why your bill depends on it
A token is the small chunk of text that an AI model actually reads and writes. It is not a word and it is not a letter. It sits somewhere in between, and every major AI vendor counts usage in tokens and charges you for them.
I run a few small sites and automations out of Singapore, and the first time I got a surprising API invoice, the cause was simple: I had been thinking in words and pages while the meter was running in tokens. If you use a chatbot on a flat monthly plan you can ignore all this. If you ever call a model through an API, build a tool on top of one, or wonder why a long conversation gets slower and pricier, tokens are the unit that explains it.
what it is
A token is a piece of text that a model treats as one unit. Before a language model sees your message, a separate step called a tokenizer chops the text into pieces and turns each piece into a number. The model only ever works with those numbers.
Pieces can be:
- a whole common word, like “the” or “price”
- part of a longer word, so “tokenization” might become “token” and “ization”
- a single character or punctuation mark
- a space attached to the start of a word
OpenAI publishes a rule of thumb for English text: one token is roughly four characters, and 100 tokens is roughly 75 words. You can test this yourself with OpenAI’s online tokenizer, which colours each token in whatever text you paste. It is the fastest way to build a feel for the unit.
Two caveats from my own testing. First, every vendor uses its own tokenizer, so the same sentence can produce different counts on different models. Second, the four characters rule is for English. Other languages, code, and unusual formatting often break into more tokens per word. Chinese, Malay and Tamil text, all common here in Singapore, usually costs more tokens for the same meaning than English does, and that matters if you build for local audiences.
how it works
There are three parts to understand: how text becomes tokens, what the model does with them, and how the bill gets calculated.
from text to tokens
The tokenizer is trained on a large body of text to find the chunks that appear most often. Frequent chunks get their own token. Rare words get split into smaller known pieces. That is why an everyday word costs one token while a long product name or a typo may cost four or five.
what the model does with tokens
A model reads your input tokens, then generates its answer one token at a time. Each new token is predicted from everything before it. That one-at-a-time process is also why output feels like it streams onto the screen. If you want to see how the choice of each next token gets made, my piece on temperature, top-p and sampling covers it, and speculative decoding covers one trick vendors use to make that process faster.
the context window
Every model has a context window, which is the maximum number of tokens it can handle in one request. Importantly, that limit covers your input and the model’s output together. Your system instructions, the conversation so far, any documents you paste in, and the reply all share the same budget. When a chat gets long, the whole history is usually sent again with each new message, so token use grows as the conversation grows.
how the bill is calculated
API vendors price per token, usually quoted per million tokens, and they almost always charge different rates for input and output. Output tokens typically cost more than input tokens. You can see current rates on Anthropic’s pricing page and OpenAI’s API pricing page. I am deliberately not quoting numbers here, because they change often and an old figure in an article is worse than none.
Here is an illustrative calculation. The prices are made up for the example, so check the real ones before you budget.
- assumed input price: $3 per million tokens
- assumed output price: $15 per million tokens
- one request sends 2,000 input tokens (instructions plus a pasted document) and gets 500 output tokens back
- input cost: 2,000 x $3 / 1,000,000 = $0.006
- output cost: 500 x $15 / 1,000,000 = $0.0075
- total per request: $0.0135
That looks like nothing. Run it 10,000 times a day and it is $135 a day. The per-request cost is tiny, and the volume is what hurts. This is exactly how my first surprise invoice happened.
Most vendors also return the token counts in every API response, so you can log what each call really cost instead of guessing. Anthropic’s documentation on token counting shows how to count tokens before you send a request.
why it matters
it is the unit your cost is measured in
If you are building anything on an API, cost per request equals tokens in times the input rate plus tokens out times the output rate. Once you can estimate tokens, you can estimate a monthly bill before launch rather than after. That is the difference between a pricing decision and a nasty surprise.
it shapes how you write prompts
Long instructions, pasted documents and repeated examples all cost input tokens on every single call. A 3,000 token system prompt that looks harmless becomes a real line item at volume. Trimming it, or using the caching features some vendors offer for repeated prefixes, is often the cheapest optimisation available. Before you reach for something heavier, read when fine-tuning beats prompting, because one of the trade-offs there is exactly this: a tuned model can need shorter prompts.
it limits what the model can see
Because the context window is measured in tokens, it decides how much of a document, codebase or conversation fits in one go. When something does not fit, you have to cut it, summarise it or retrieve only the relevant pieces. That last approach is the idea behind retrieval systems, and when they go wrong the answers can be confidently off, which I get into in why your RAG answers are wrong.
it affects speed as well as cost
More input tokens take longer to process, and more output tokens take longer to generate, since output is produced one token at a time. A response that is 2,000 tokens long will always arrive more slowly than one that is 200, on the same model. If your product feels sluggish, the token count is the first thing I would check.
common misconceptions
“a token is a word”
It is not. Short common words are often one token, but longer or rarer words split up, and punctuation and spaces count too. The four characters per token rule is only an average for English. Paste your own text into a tokenizer instead of assuming.
“I only pay for what I type”
You pay for input and output, and the input includes everything the model is sent: system prompt, history, attached documents and tool definitions. In a long chat the history alone can dwarf your latest message. Output tokens also usually cost more per token than input, so a chatty model costs more than a terse one.
“all models count tokens the same way”
Different vendors and even different model generations use different tokenizers. The same paragraph can be a different token count on two models, so a price per million tokens is not directly comparable across vendors until you measure your own text on each. When I compare options, I run the same real sample through each and compare cost per task, not the headline rate.
“a bigger context window means I should fill it”
A large window is a ceiling, not a target. You pay for every token you send, and stuffing a window with loosely relevant material tends to raise cost and slow responses without guaranteeing a better answer. Sending the right few thousand tokens usually beats sending everything.
where to go from here
Once tokens make sense, the next questions are usually about money and about model choice. These are the pieces I would read next:
- committed spend deals and the break-even point: how prepaying for usage changes your effective rate, and when it is worth it
- what a per-seat price hides: why a flat seat price and a per-token price behave so differently once usage grows
- the best open-source LLMs you can self-host in 2026: the option where you stop paying per token and start paying for hardware
- building an eval set from support tickets: how to measure whether a cheaper or shorter setup still gives good answers
If your prompts include customer data, remember that those tokens are being sent to a third party. The team at The Privacy Wire writes about what that means in practice, and it is worth a read before you pipe anything sensitive into an API. For everything else I have published on this topic, the blog index has the full list.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-02.