Russian-language text processed by the American neural network Claude can cost the user significantly more than the same text in GPT. A study by Donseo showed that the same Russian text in Claude Opus 5 occupied approximately 3.9 times more tokens than in GPT-5.5. With the same token cost, this leads to an almost fourfold difference in expenses.
The difference becomes noticeable when converted to real tariffs. The author of the study compared the median prices of Russian services and found about 0.53 rubles per 1000 Russian characters when using Claude Opus 5 versus approximately 0.14 rubles for GPT-5.5. For a single request, the difference is only a few kopecks, but for chatbots, support services, and other services with millions of messages, additional costs can become significant.
The reason is not a separate charge for the Russian language, but the operation of the tokenizer. Before processing, the neural network breaks the text into tokens - individual elements that the model then works with. Different models do this differently: GPT is able to combine several Russian characters into one token, while Claude Opus 5 in the study often broke Cyrillic text into a significantly larger number of tokens.
As a result, a phrase with the same meaning can have a completely different volume for different neural networks. The more tokens a model needs to process a request, the higher the cost when paying by the number of tokens.
The author of the study compared 11 families of neural networks. For GPT, Gemini, Grok, and GLM, Russian text required approximately 18–23% more tokens than analogous English. For DeepSeek, the difference was about 38%, for Qwen - 52%, and for Kimi - 86%.
The largest gap was shown by Claude Opus 5: the Russian version of the text occupied almost three times more tokens than the English one. When compared specifically with GPT-5.5, the difference for the same Russian text reached approximately 3.9 times.
Russian models showed the opposite result. In YandexGPT and GigaChat, Russian text occupied even fewer tokens than its English translation. This means that when choosing a model for a Russian-language service, a single price per million tokens is not enough: it is also necessary to consider the number of tokens that the model spends on processing real text.
For small requests, this difference is almost imperceptible, but with large volumes of data, it directly affects the cost of the service. Therefore, it makes sense to compare neural networks for Russian-language products not only by the stated tariff, but also by the actual consumption of tokens in Russian.