
Stop Wasting Tokens on Prompts
Every prompt token is billed on every single call. A system prompt with 2,000 tokens costs you 2,000 tokens per request — multiply that across thousands of calls and it dwarfs the output cost.

Three quick wins 🏆
- Cut boilerplate your model ignores (rules it already follows)
- Move rarely-needed instructions out of the system prompt
- Trim chat history to the last N messages that matter
| Prompt change | Token saving | Quality impact |
|---|---|---|
| Drop redundant rules | 20–30% | None |
| Shorten history | 10–40% | Watch the first turns |
| Compress examples | 5–15% | Minimal |
Measure your prompt size 📏
Log token usage per call and review which requests carry the biggest prompts. A quick weekly review of the top ten usually surfaces obvious waste.
Frequently asked questions ❓
How do I see my prompt token usage?
The Usage page breaks spend down by tokens in and out, so you can spot where prompts dominate.
Will shorter prompts hurt quality?
Usually not — most long prompts contain instructions the model already obeys by default.
What about long documents?
For big documents, retrieve only the relevant sections instead of sending the whole file every time.
Comments (0)