
Vision Models: When They Earn Their Keep
Vision models can read screenshots, diagrams, documents and photos. They are surprisingly good — but they cost more, so the trick is knowing when they pay off.

Great use cases ✅
- Extracting tables from screenshots or scanned PDFs
- Validating UI mockups against a spec
- Reading labels, receipts or handwritten notes
- Describing images for accessibility
Skip them for these ❌
Do not send images just because you can. If the information is already in text or structured data, a text model is cheaper and just as accurate.
Mind the token math 🧮
| Factor | Effect on cost |
|---|---|
| Image resolution | Higher resolution = more tokens |
| Many images | Cost scales with each image |
| Text overlay | OCR-heavy images behave like long text |
Downscale images to the smallest size that still contains the information you need, and crop out irrelevant borders before sending.
Frequently asked questions ❓
What image formats are supported?
Common web formats such as PNG, JPEG and WebP work out of the box.
Are vision models slower?
Yes, processing pixels takes longer than text alone, so budget for higher latency.
Can I resize images before sending?
Definitely — sending a 4000px screenshot when a 1000px one suffices just burns tokens.
Comments (0)