Code & data · tokensToken counter
How many tokens a text takes in ChatGPT and open models — and why Russian text costs more than English. Counted in your browser; the text goes nowhere.
Count
pick a model vocabularyRussian vs English
our test: Universal Declaration of Human Rights, art. 1–3Closed tokenizers
no made-up coefficientsClaude · Anthropic
The tokenizer is not published. An exact count comes only from the provider's API: Anthropic's free count_tokens method. According to Anthropic's documentation, models from Claude Opus 4.7 onward use a newer tokenizer, and the same text takes roughly 30 percent more tokens than on earlier models.
Gemini · Google
The tokenizer is not published. An exact count comes only from the provider's API: Google's countTokens method.
≈ 44 by Google's rule “1 token ≈ 4 characters” — undercounts Russian
YandexGPT · Yandex
The tokenizer is not published. An exact count comes only from the provider's API: Yandex's Tokenizer method.
About tokens
AI models read text neither by letters nor by words but by tokens — pieces of words from the model's vocabulary. The context limit and the price of a request depend on the number of tokens. Every model has its own vocabulary, so the same text takes a different number of tokens, and Russian usually more than English: the vocabularies are built mostly from English text.
Token
Unito200k
GPT-4o and newercl100k
GPT-4 and GPT-3.5Russian costs more
InequalityClosed vocabularies
Claude, GeminiInvisible characters
Extra tokensFrequently asked questions
Updated
Counting runs in your browser. We do not show prices: they change, and the bill also includes the answer, reasoning and service tokens.