Code & data · tokensToken counter

How many tokens a text takes in ChatGPT and open models — and why Russian text costs more than English. Counted in your browser; the text goes nowhere.

Counted in your browser · vocabularies load from bringer.ru, not from third-party servers
01Count02Russian vs English03Closed04About tokens05Questions06Related
01

Count

pick a model vocabulary
ClosedClaude not publishedGemini not publishedYandexGPT not published
Text
175 characters · in-browser limit — several megabytes
Tokens
—
Characters
175
Words
29
Characters per token
—
loading the vocabulary…
Context window— of 128,000 used · —
02

Russian vs English

our test: Universal Declaration of Human Rights, art. 1–3
o200k · GPT-4o and newer1.43×
English146
Russian209
cl100k · GPT-4 and GPT-3.52.70×
English146
Russian394
For comparison: Petrov et al., NeurIPS 2023 — cl100k on Russian ≈ 2.49× English (FLORES-200).Your text: — characters per token
03

Closed tokenizers

no made-up coefficients

Claude · Anthropic

The tokenizer is not published. An exact count comes only from the provider's API: Anthropic's free count_tokens method. According to Anthropic's documentation, models from Claude Opus 4.7 onward use a newer tokenizer, and the same text takes roughly 30 percent more tokens than on earlier models.

Gemini · Google

The tokenizer is not published. An exact count comes only from the provider's API: Google's countTokens method.

≈ 44 by Google's rule “1 token ≈ 4 characters” — undercounts Russian

YandexGPT · Yandex

The tokenizer is not published. An exact count comes only from the provider's API: Yandex's Tokenizer method.

04

About tokens

AI models read text neither by letters nor by words but by tokens — pieces of words from the model's vocabulary. The context limit and the price of a request depend on the number of tokens. Every model has its own vocabulary, so the same text takes a different number of tokens, and Russian usually more than English: the vocabularies are built mostly from English text.

Token

Unit
to·ken·izer
A piece of text from the model's vocabulary: a common word whole, a rare one in parts. A leading space is usually part of the token.

o200k

GPT-4o and newer
привет = 2
OpenAI's vocabulary from GPT-4o onward. Russian text costs almost half as much in it as in the GPT-4 vocabulary.

cl100k

GPT-4 and GPT-3.5
привет = 4
OpenAI's older vocabulary. Russian text takes 2.5–2.7 times as many tokens as English in it.

Russian costs more

Inequality
1.4× · 2.7×
In our test Russian text took 1.4 times as many tokens as English in o200k and 2.7 times in cl100k.

Closed vocabularies

Claude, Gemini
≈
Anthropic and Google do not publish their tokenizers. Only their APIs give an exact count.

Invisible characters

Extra tokens
U+202F = 1–2
Invisible characters cost tokens too. Cleaning the text before sending removes them.
05

Frequently asked questions

A piece of text a model works with: a common word whole, a rare one in several parts, and a leading space is usually part of the token. Every model has its own vocabulary of such pieces, so the same text takes a different number of tokens. Both the context limit — how much text the model sees at once — and API prices are counted in tokens.

Updated

Counting runs in your browser. We do not show prices: they change, and the bill also includes the answer, reasoning and service tokens.

«» added to favorites