
The same text costs 50 to 84% more in Serbian than in English across 8 AI models, and another 5 to 24% more in Cyrillic. We measured it on 2026 news.
When a company uses GPT, Gemini or any other language model through an API, it does not pay per word. It pays per token. A token is a piece of text from the model's vocabulary: sometimes a whole word, sometimes a syllable or a single letter. The more pieces a text falls into, the bigger the bill, the slower the answer and the less text the model can hold in memory at once.
We wanted to know how many pieces Serbian text falls into, and whether the script matters. Serbian is written in two scripts, Cyrillic and Latin, and the same sentence can be written in either. The answer: it matters. In every one of the eight models we measured, the same Serbian text in Cyrillic uses more tokens than in Latin script.
How we measured
We took 136 articles from Serbia's public broadcaster RTS, published between April and October 2026: 2,157 paragraphs in total. The texts are written in Cyrillic, and we transliterated them letter by letter into Latin script. That gave us two versions of exactly the same text, where the only difference is the script.
To compare with English and other languages we used Microsoft's NTREX-128 set of professional translations, where the same 1,997 sentences exist in English, Serbian, Croatian, Bosnian, Slovenian, Macedonian, Bulgarian, Russian and Ukrainian.
We ran every text through the tokenizers of GPT-5 (OpenAI), Gemini (Google, through the open Gemma 4 model, which Google's June 2026 report says uses the same tokenizer as Gemini), Llama 4 (Meta), Mistral Medium 3.5, DeepSeek V4.1, Qwen3.8 (Alibaba), Kimi K3 (Moonshot) and GLM-5.3 (Z.ai).
Research published this year showed that tokenizers charge more for Slavic languages, Ukrainian in particular. We found no measurement for Serbian Latin and Cyrillic on the newest models and on Serbian texts from 2026, so we did it ourselves.
What we found
Serbian costs more than English in every model. On GPT-5 the same content takes 100 tokens in English, 153 in Serbian Latin and 180 in Serbian Cyrillic. Across all eight models, Latin script costs 50 to 84% more than English, and Cyrillic 65 to 129% more.
| Language (GPT-5) | Tokens for the same text |
|---|---|
| English | 1.00× |
| Russian | 1.47× |
| Bosnian | 1.49× |
| Croatian | 1.50× |
| Slovenian | 1.52× |
| Serbian, Latin script | 1.53× |
| Bulgarian | 1.77× |
| Ukrainian | 1.79× |
| Serbian, Cyrillic | 1.80× |
| Macedonian | 1.81× |
Cyrillic costs more than Latin script in all eight models. On the RTS news the gap is largest on Kimi K3, 24%, then GPT-5, 17%, and Google's Gemini, 15%. Qwen3.8 and GLM-5.3 charge 14% more for Cyrillic, Mistral 7%, Llama 4 and DeepSeek 5%. On GPT-5 a Cyrillic paragraph costs more than its Latin twin in 99% of cases. Microsoft's translation set gave almost the same numbers, so the result does not depend on the choice of texts.
| Model | Cyrillic vs Latin, same text |
|---|---|
| Kimi K3 | +24% |
| GPT-5 | +17% |
| Gemini | +15% |
| Qwen3.8 | +14% |
| GLM-5.3 | +14% |
| Mistral Medium 3.5 | +7% |
| Llama 4 | +5% |
| DeepSeek V4.1 | +5% |
Russian does better than Serbian Cyrillic. Russian, also written in Cyrillic, uses fewer tokens than Serbian Cyrillic in all eight models. On GPT-5 Russian takes 1.47 tokens for every English token, less even than Serbian in Latin script. So the problem is not only the script.
Why it happens
A tokenizer builds its vocabulary by looking for the most frequent pieces of words in a huge amount of text. What appears often gets its own token. What is rare stays broken into small pieces.
We looked at the vocabulary GPT-5 uses. Of about 200,000 tokens, more than 14,000 contain Cyrillic letters. Only 66 contain at least one of the letters ђ, ћ, џ, љ, њ or ј, which exist in Serbian but not in Russian. Russian letters that Serbian does not use (ы, э, ъ and ё) appear in 1,693 tokens, more than 25 times as often. Other models look similar: between 12 and 170 tokens with Serbian letters, between 193 and 2,324 with Russian ones.
You can see the result on ordinary words. The Russian word „люди“ (people) is one token. The Serbian word „људи“, one letter different, is cut into two: „љ“ and „уди“. The same happens with „још“ (still) and „његов“ (his). In Latin script, „ljudi“, „još“ and „njegov“ are one token each, because Latin writes those sounds with letter pairs the model knows well.
The hidden letter that makes text more expensive
While measuring we ran into a problem you cannot see. In the professional Cyrillic translation from Microsoft's set we found 8,551 words with a Latin letter mixed in: „a“, „e“, „o“ or „j“. They look exactly like their Cyrillic twins, but to the model they are a different script. The word „који“ (which) in clean Cyrillic is one token. With a Latin „o“ and „j“ inside, the same word becomes three. Because of such letters, that text used 19% more tokens on GPT-5 than after cleaning. The RTS news has such words too, but far fewer: about fifty in the whole sample.
What it means in money and time
For someone who asks a chatbot a question now and then, the difference goes unnoticed. For a company processing tens of thousands of messages a day through an API, it shows up at the end of the month. If a job on GPT-5 costs €100 in English, the same content costs about €153 in Serbian Latin and about €180 in Cyrillic.
The same goes for the model's memory. Where a 100-page document fits in English, the same document fits only up to page 65 in Serbian Latin and page 56 in Cyrillic. With long contracts, policies or technical documentation, the model "forgets" the beginning sooner.
Since a model writes its answer token by token, a longer text also means a slower answer. At the same speed, an answer that arrives in 10 seconds in English takes about 15 seconds in Latin script and about 18 in Cyrillic.
Nobody did this on purpose. Tokenizers are built on the texts there are most of, and Serbian, Cyrillic in particular, is clearly rare in that data. The effect is still the same as a tax: whoever writes in Serbian pays more, and whoever writes in Cyrillic pays even more.
Developers and chatbots
We measured code too. A simple JavaScript function with English variable names and comments is 142 tokens on GPT-5. The same function with Serbian names and comments is 205, or 44% more.
A separate cost is the instructions sent to the model with every request, the system prompt. A short instruction for a customer-support chatbot is 88 tokens in English, 123 in Serbian Latin and 143 in Cyrillic. That difference is paid on every message from every user.
Two myths
Myth one: Latin script without diacritics saves tokens. On GPT-5 and Google's tokenizer there is no saving; text without diacritics is even slightly more expensive. On the other models the saving is at most 4%.
Myth two: one of the region's languages is "smarter" for AI. Serbian Latin, Croatian, Bosnian and Slovenian cost almost the same on GPT-5: between 1.49 and 1.53 tokens per English token. The difference only appears when the script changes.
Local models don't fix it
The Serbian models published on Hugging Face in 2026 that we checked were fine-tuned from the Chinese model Qwen. They kept its tokenizer, so they use as many tokens on the same text as the original. A model can learn to write better Serbian, but if the vocabulary stays the same, so does the bill.
What companies can do today
- Clean hidden Latin letters out of texts before sending them to the model. It is a simple check a program runs in a fraction of a second, and for some texts it cuts the cost by a fifth.
- Ask for the answer in Latin script and show it in Cyrillic. A company that must answer users in Cyrillic can ask the model to answer in Latin and transliterate before display. On GPT-5 that cuts the tokens in the answer by about 15%. Serbian transliteration from Latin to Cyrillic is almost entirely mechanical; the rare exceptions, like „injekcija“ or „nadživeti“, where „nj“ and „dž“ are not single letters, and foreign names, are handled with a short exception list.
- Measure models on your own texts. In English, the most and least efficient of the eight models differ by 4%. In Serbian Latin the gap is 22%, in Cyrillic 36%.
Before choosing a model for a product in Serbian, or any smaller language, measure how many tokens it uses on that company's own texts. The gap between models is much larger than in English, and it is paid every day. If you are building an AI product or automation in a language other than English, we run this measurement before the model is chosen.
Limits of the study
We measured token counts, not answer quality. Whether models also make more mistakes in Cyrillic is a topic for the next study. Anthropic's Claude is not included because its tokenizer is not public. For Gemini we used the Gemma 4 tokenizer, which Google says is the same as Gemini's. The euro amounts illustrate the ratio and are not a price list: the real bill depends on the model, its price and the mix of input and output tokens. Comparisons with other languages also depend on the translation; the Latin vs Cyrillic comparison does not have that problem, because the text is identical. Montenegrin was not measured separately.
The full data, measurement code and charts are available on request, so anyone can repeat the measurement. Write to us. The original study in Serbian, with charts, is here.
Sources
- Gemma Team, Google DeepMind (June 2026), Gemma 4 Technical Report
- Ovcharov (May 2026), The Tokenizer Tax Across 25 European Languages
- Dobrovolskyi (July 2026), Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems
- RTS news, April to October 2026
- Microsoft, NTREX-128, professional translation set
Frequently asked questions
Why does AI cost more in Serbian than in English?
Language models do not charge per word but per token, a piece of text from the model's vocabulary. The vocabulary is built from the texts that are most common, and Serbian is rare in them. Serbian words are cut into more pieces, so the same content uses more tokens: on GPT-5 about 53% more in Latin script and 80% more in Cyrillic than in English.
Is Cyrillic more expensive than Latin script in every model?
Yes. In all eight models we measured, the same text in Cyrillic uses more tokens than in Latin script: Kimi K3 24% more, GPT-5 17%, Gemini 15%, Qwen3.8 and GLM-5.3 14%, Mistral 7%, Llama 4 and DeepSeek 5%. On GPT-5 a Cyrillic paragraph costs more than its Latin twin in 99% of cases.
How can a company lower the cost of AI in Serbian?
Clean texts of Latin letters mixed into Cyrillic before sending them to the model, ask the model to answer in Latin script and transliterate the answer to Cyrillic before showing it (about 15% fewer tokens on GPT-5), and measure models on the company's own texts before choosing one, because the gap between models is up to 36% in Serbian and only 4% in English.


