Tokenizers and why some languages cost more
If you’ve ever run the same prompt through a translator and then through an LLM API in two different languages, you’ve probably noticed the token count doesn’t match the word count. English input feels cheap. Input in Japanese, Thai, or Arabic often costs noticeably more for what feels like the same amount of content. This isn’t a pricing gimmick. It comes directly from how tokenizers are built and trained, and it’s worth understanding if you’re shipping a product that touches more than one language.
What a tokenizer actually does
Before any text reaches a language model, it gets converted into a sequence of tokens, numeric IDs that map to chunks of text the model was trained to recognize. Most current LLMs use some flavor of byte pair encoding (BPE) or a close relative like SentencePiece’s unigram model. The idea is simple: start with individual characters or bytes, then repeatedly merge the most frequently co-occurring pairs into single tokens, building up a fixed vocabulary of maybe 50,000 to 200,000 entries.
The word “tokenization” gets applied here, and it matters that this vocabulary isn’t hand designed. It’s learned from a training corpus, usually a huge scrape of web text, code, and books. Whatever character sequences show up most often in that corpus get merged into their own tokens first. Common English words like “the,” “and,” or “tion” end up as single tokens early in training because they’re everywhere in the source data. Rare sequences stay broken into smaller pieces, sometimes down to individual characters or bytes.
That last part is the whole story. Token efficiency isn’t a property of a language. It’s a property of how well a language’s common substrings are represented in whatever corpus trained the tokenizer.
Why English trains cheap
The corpora used to train the big public tokenizers are dominated by English and other Latin-script, high-resource languages. Common Crawl, GitHub, Wikipedia dumps, they all skew heavily toward English content, both in raw volume and in how much gets kept after quality filtering. When BPE runs its merge algorithm over that data, it spends most of its limited vocabulary budget encoding English morphology efficiently: word stems, common suffixes, whole common words, even frequent multi-word phrases like “as well as” sometimes collapse toward fewer tokens.
A language that makes up a smaller share of the training data doesn’t get that treatment. Its common word-forming patterns are rarer in the corpus, so the merge algorithm never prioritizes them into single tokens. The vocabulary slots get spent elsewhere. The practical result is that a sentence in a lower-resource language ends up split into more, smaller tokens to represent the same meaning, because the tokenizer never learned efficient shortcuts for it.
This is why the fix isn’t really about the language itself, Finnish and Hungarian have complex morphology and still tokenize reasonably well when they’re well represented in training data. It’s about representation in the specific corpus that trained that specific tokenizer.
Scripts change the math before you even hit the vocabulary
There’s a second, more mechanical factor layered on top of corpus representation: character encoding. Most modern tokenizers, including the byte-level BPE used by several major model families, operate on raw UTF-8 bytes rather than Unicode codepoints directly. That matters because UTF-8 doesn’t use the same number of bytes for every script.
Basic Latin characters (the ones used in English, and most of the alphabet-based languages that share it) encode as a single byte each in UTF-8. Extended Latin, Cyrillic, Greek, Hebrew, and Arabic characters typically take two bytes. Characters in CJK scripts (Chinese, Japanese, Korean) generally take three bytes. So before the BPE merge rules even get a say, a Japanese sentence is already starting from more raw bytes per character than an English sentence of equivalent length. If the tokenizer’s vocabulary hasn’t learned good multi-byte merges for that script, those extra bytes stay closer to their raw, unmerged form, and each character can end up costing more than one token on its own.
Scripts that don’t use whitespace to separate words add a third wrinkle. English tokenizers lean heavily on the space character as a natural word boundary, which is part of why common English words merge cleanly. Thai, Burmese, and to a lesser extent Chinese and Japanese don’t segment words with spaces at all. A tokenizer trained mostly on space-delimited languages doesn’t get that same free signal for where one semantic unit ends and the next begins, so it has less to work with when deciding what to merge.
Put those three factors together, thinner training representation, more bytes per character, and no whitespace word boundaries, and you get compounding costs for some languages rather than a single clean explanation.
Where it actually bites you in production
This stops being trivia the moment you’re paying per token or budgeting a context window. A few concrete places it shows up if you’re building on top of these APIs:
Input and output cost. API pricing is quoted per token, not per character or per word. If a support ticket in one language takes meaningfully more tokens to represent than the same ticket translated to English, your per-request cost for that market is higher even though the customer typed the same amount of information. This is easy to miss if you built your cost model in English and only later added multilingual support.
Context window budget. Context limits are also token limits. A document that fits comfortably in a 128k-token context window in English can consume noticeably more of that budget when it’s the same document in a language with worse tokenization efficiency. For RAG pipelines, this means your chunk size in tokens doesn’t map to a consistent amount of actual content across languages. A chunking strategy tuned on English source documents can quietly produce smaller, less coherent chunks when applied to other languages, which hurts retrieval quality on top of costing more.
Output truncation. If you set a max output token limit expecting a certain amount of response text, a model generating in a token-hungry language can hit that ceiling and get cut off mid-thought sooner than the equivalent English response would. This is a subtle bug source: the model isn’t failing to reason, it’s just running out of budget faster.
Latency. Generation happens token by token. More tokens needed to say the same thing generally means more decoding steps, which means slower responses for users writing in less token-efficient languages, all else equal.
What you can actually do about it
There’s no trick that erases this cost difference, since it’s baked into how the vocabulary was trained, but there are ways to work around it rather than get surprised by it.
Measure it for your own workload instead of assuming. Tokenizer libraries like OpenAI’s tiktoken are public and let you count tokens for real strings before you ship. Run your actual prompts and expected outputs, in the actual languages your users write in, and look at the token counts directly rather than guessing from character counts.
Build cost models per language, not per feature. If you’re pricing a product or setting internal budgets, a flat “average cost per request” number will undercount your cost on token-heavy languages and overcount on token-cheap ones. Break it out.
Size RAG chunks in tokens, not characters, and re-tune chunk size per language if you’re serving a genuinely multilingual corpus. A chunk boundary that makes sense in English may split a sentence awkwardly in a script that tokenizes less efficiently.
Leave headroom on output limits for languages you know are less efficient, rather than using one global max-tokens setting tuned against English testing.
None of this changes the underlying economics. It just means you’re not finding out about it from a support ticket or a surprise invoice.
If you want more breakdowns like this on how the tools you’re actually shipping with work under the hood, you’ll find them on the AI Tool Gazette home page.