Tokens, Not Words: How Tokenization Affects Your Bill

Tokens, Not Words: How Tokenization Affects Your Bill

Words are not tokens 🔤

Models do not read words, they read tokens. A token is a chunk of text: often a whole word, sometimes part of a word, punctuation or whitespace. In English one token is roughly four characters, or about three quarters of a word 📏.

How text becomes tokens
The same idea can cost a different number of tokens in each language

Why the same text costs more in some languages 🌍

Tokenizers are trained mostly on English text, so English is usually the most efficient language. Other languages and code can need noticeably more tokens for the same amount of information.

Content typeApprox. tokens per wordNotes
English prose1.3Most efficient
European languages1.5 - 1.8More sub-word splits
Chinese / Japanese~1 per characterDepends on the tokenizer
Source code1.5 - 2.5Symbols and indentation count

Estimating tokens before you send 📐

You can estimate cost before a request by counting tokens with a tokenizer that matches the model family you use. A rough fallback is to count characters and divide by four, then add a small safety margin.

  • Count the prompt and the expected reply, not just the prompt
  • Long system prompts are charged on every single call
  • Repeated boilerplate adds up fast across thousands of requests
  • Measure real usage after launch instead of guessing forever

Designing prompts with tokens in mind 🧠

Trim instructions that never change the answer, move fixed reference text into a cached or retrieved block where supported, and keep few-shot examples short. Small prompt savings multiply across every request you make.

Frequently asked questions ❓

Do I pay for both the prompt and the reply?

Yes. Input tokens and output tokens are both billed, and output tokens usually cost more per token.

Is one word always one token?

No. Short common words are often a single token, while long or unusual words split into several.

Why did my cost change without changing my prompt?

Different models use different tokenizers, and a model upgrade can change how the same text is tokenized.

Should I translate my prompts to English?

Only if it does not hurt quality. English is often more token-efficient, but your users care about the answer, not the prompt language.

Share your thoughts
English

Did this post help you? Leave a comment below — we read every single one and reply to as many as we can.

Comments are moderated and published after a quick review.

Español

¿Te ayudó esta publicación? Deja un comentario abajo: leemos todos y respondemos a la mayoría.

Los comentarios se moderan y se publican tras una revisión rápida.

中文

这篇文章对您有帮助吗?请在下方留言——我们会阅读每一条评论并尽可能回复。

评论需经过审核后才会显示。

Deutsch

Hat dir dieser Beitrag geholfen? Hinterlasse unten einen Kommentar – wir lesen jeden einzelnen und antworten so vielen wie möglich.

Kommentare werden moderiert und nach einer kurzen Prüfung veröffentlicht.

Français

Cet article vous a-t-il aidé ? Laissez un commentaire ci-dessous – nous les lisons tous et répondons au plus grand nombre.

Les commentaires sont modérés et publiés après une vérification rapide.

Português

Este artigo ajudou você? Deixe um comentário abaixo – lemos todos e respondemos à maioria.

Os comentários são moderados e publicados após uma revisão rápida.

Nederlands

Heeft dit bericht je geholpen? Laat hieronder een reactie achter – we lezen ze allemaal en beantwoorden zoveel mogelijk.

Reacties worden gemodereerd en na een korte controle gepubliceerd.

Polski

Czy ten post Ci pomógł? Zostaw komentarz poniżej – czytamy każdy i odpowiadamy na tyle, na ile możemy.

Komentarze są moderowane i publikowane po szybkiej weryfikacji.

Русский

Эта статья помогла вам? Оставьте комментарий ниже — мы читаем каждый и отвечаем на большинство.

Комментарии модерируются и публикуются после быстрой проверки.

العربية

هل ساعدك هذا المقال؟ اترك تعليقًا أدناه — نقرأ كل تعليق ونرد على أكبر عدد ممكن.

تتم مراجعة التعليقات وتنشر بعد التحقق السريع.

日本語

この記事は役に立ちましたか?下のコメント欄にぜひ投稿してください。すべて読んで、できる限り返信します。

コメントはモデレーション後に公開されます。

한국어

이 글이 도움이 되었나요? 아래에 댓글을 남겨주세요 — 모든 댓글을 읽고 최대한 많이 답변합니다.

댓글은 검토 후 게시됩니다.

ไทย

บทความนี้ช่วยคุณได้ไหม? แสดงความคิดเห็นด้านล่าง — เราอ่านทุกความเห็นและตอบกลับให้มากที่สุด

ความคิดเห็นจะถูกตรวจสอบและเผยแพร่หลังจากการพิจารณา

עברית

האם המאמר עזר לכם? השאירו תגובה למטה — אנחנו קוראים כל תגובה ועונים לרובן.

תגובות עוברות ניהול ומתפרסמות לאחר בדיקה מהירה.

Tiếng Việt

Bài viết này có giúp bạn không? Hãy để lại bình luận bên dưới — chúng tôi đọc từng bình luận và phản hồi nhiều nhất có thể.

Bình luận được kiểm duyệt và xuất hiện sau khi xét duyệt nhanh.

Indonesia

Apakah artikel ini membantu Anda? Tinggalkan komentar di bawah — kami membaca semuanya dan membalas sebanyak mungkin.

Komentar dimoderasi dan dipublikasikan setelah pemeriksaan cepat.

Қазақша

Бұл мақала сізге көмектесті ме? Төменде пікір қалдырыңыз — біз әрқайсысын оқып, мүмкіндігінше көбіне жауап береміз.

Пікірлер модерациядан өтіп, тексерілгеннен кейін жарияланады.

Comments (0)

Demo preview: the sample comments below are seeded for layout preview. Newly submitted comments are moderated before they appear.
No comments yet. Be the first to share your thoughts!