
Tokens, Not Words: How Tokenization Affects Your Bill
Models do not read words, they read tokens. A token is a chunk of text: often a whole word, sometimes part of a word, punctuation or whitespace. In English one token is roughly four characters, or about three quarters of a word 📏.

Why the same text costs more in some languages 🌍
Tokenizers are trained mostly on English text, so English is usually the most efficient language. Other languages and code can need noticeably more tokens for the same amount of information.
| Content type | Approx. tokens per word | Notes |
|---|---|---|
| English prose | 1.3 | Most efficient |
| European languages | 1.5 - 1.8 | More sub-word splits |
| Chinese / Japanese | ~1 per character | Depends on the tokenizer |
| Source code | 1.5 - 2.5 | Symbols and indentation count |
Estimating tokens before you send 📐
You can estimate cost before a request by counting tokens with a tokenizer that matches the model family you use. A rough fallback is to count characters and divide by four, then add a small safety margin.
- Count the prompt and the expected reply, not just the prompt
- Long system prompts are charged on every single call
- Repeated boilerplate adds up fast across thousands of requests
- Measure real usage after launch instead of guessing forever
Designing prompts with tokens in mind 🧠
Trim instructions that never change the answer, move fixed reference text into a cached or retrieved block where supported, and keep few-shot examples short. Small prompt savings multiply across every request you make.
Frequently asked questions ❓
Do I pay for both the prompt and the reply?
Yes. Input tokens and output tokens are both billed, and output tokens usually cost more per token.
Is one word always one token?
No. Short common words are often a single token, while long or unusual words split into several.
Why did my cost change without changing my prompt?
Different models use different tokenizers, and a model upgrade can change how the same text is tokenized.
Should I translate my prompts to English?
Only if it does not hurt quality. English is often more token-efficient, but your users care about the answer, not the prompt language.
Comments (0)