
Audit Your Spend by Model and Token Type
A single monthly number tells you how much you spent, never why. The fastest way to control cost is to break usage into a few dimensions and look at the biggest bars first.

The four useful cuts 🔪
| Dimension | Example | Question it answers |
|---|---|---|
| By model | fast vs large | Which model costs most? |
| By feature | chat vs batch jobs | Which product line drives spend? |
| By token type | input vs output | Are long replies expensive? |
| By customer | free vs paid tiers | Does revenue cover usage? |
A monthly routine 🗓️
- Export usage and group it by model and by day
- Compare this month with the previous month, not with zero
- Flag any single model that grew more than the rest
- Check that your cheapest adequate model handles the bulk of traffic
Act on the biggest bar 🎯
Usually one model or one feature dominates the bill. Fix that first: route simple jobs to a smaller model, shorten repeated system prompts, or cap reply length where users rarely read the end.
Frequently asked questions ❓
How often should I audit usage?
Monthly is enough for most teams, plus an extra look after any big launch or traffic spike.
What if output tokens dominate?
Constrain reply length, ask for concise formats and check whether users actually read long answers.
Should every request use the cheapest model?
No. Use the cheapest model that meets the quality bar for each job, and keep a larger model for the tasks that need it.
Do I need spreadsheets?
A simple export and a couple of grouped totals are enough to spot the driver.
Comments (0)