
AI models bill per token, and English is the cheapest language there is. The same work can cost 50 to 129% more in another language. Here is why, and how to check yours.
Every business pricing an AI feature does the same thing: reads the provider's price per million tokens, runs a quick demo in English, and multiplies. The demo is in English because that is how these tools are presented. The bill is not in English. It is in your language, and your language almost certainly costs more.
Models bill per token, and English is the floor
A large language model does not charge per word. It charges per token: a piece of text from the model's fixed vocabulary, sometimes a whole word, sometimes a syllable, sometimes a single letter. The more pieces your text is cut into, the higher the bill, the slower the answer, and the less text the model can hold in memory at once.
That vocabulary is built from the text the model was trained on, which is overwhelmingly English. English words tend to survive as single tokens. Rarer languages are chopped into more, smaller pieces. So the same meaning, in another language, quietly costs more on every single call.
We measured it, and the gap is large
We ran the same text through the tokenizers of eight current models, including GPT-5, Gemini, Llama 4, Mistral, DeepSeek, Qwen, Kimi and GLM. The full method and numbers are in our study, AI Charges a Tax on Serbian, and Cyrillic Pays More. The short version:
- The same content that takes 100 tokens in English takes about 153 in Serbian Latin script and 180 in Serbian Cyrillic on GPT-5.
- Across all eight models, Latin script costs 50 to 84% more than English, and Cyrillic 65 to 129% more.
- Non-Latin scripts pay the most, and the penalty appears in every model we tested, not just one.
Serbian is one language. The mechanism is universal: any language that was rare in the training data is split into more tokens, and non-Latin scripts — Cyrillic, Arabic, the CJK scripts — tend to pay the steepest version of the tax.
Why it matters past the invoice
The token tax is not only a line on the bill. More tokens means slower responses, because the model generates one token at a time. It means less room in the context window, so long documents in your language get cut off sooner than the same documents in English. And it means the cheapest model in English is often not the cheapest model in your language: we found the gap between models reaches 36% in Serbian and only 4% in English. The ranking changes once you leave English.
How to check yours before you commit
- Measure on your own text. Take real content in your language — your listings, your emails, your documents — and count the tokens on the exact models you are considering. The provider's English benchmark will not match.
- Compare models, not just prices. Because the gap between models widens in other languages, the right model for an English product is often the wrong one for yours.
- Design around the cost. Keep system prompts and internal reasoning in English where the user never sees them, clean mixed scripts, and for some languages generate in Latin script and transliterate back.
This is exactly the kind of thing we build into custom AI automation: the automation is designed around what it actually costs to run in your language, not around an English demo. If you are planning an AI feature and want to know what it will really cost in your market, book a technical strategy call and we will measure it with you.
Frequently asked questions
Why does AI cost more in some languages than in English?
Models bill per token, a piece of text from the model's vocabulary. That vocabulary is built from the text the model saw most, which is overwhelmingly English. Rarer languages are cut into more, smaller pieces, so the same content uses more tokens and costs more. We measured Serbian at 50 to 84% more than English in Latin script and up to 129% more in Cyrillic.
Does this affect my region specifically?
It affects every language that is under-represented in the training data, which is almost every language other than English. The exact multiplier depends on the language and the script: non-Latin scripts such as Cyrillic, Arabic and the CJK scripts tend to pay the most. The only reliable number is the one measured on your own texts and the model you actually use.
How do I lower the cost of AI in my language?
Measure candidate models on your own content before choosing one, because the gap between models is far wider in other languages than in English. Keep prompts and system text in English where the user never sees them, clean mixed scripts, and for some languages ask for output in Latin script and transliterate it back. A studio that builds the automation around the token cost, not around an English demo, saves you the difference on every call.

