Start with a fast general-purpose chat model for prototypes, internal tools and high-volume tasks. It keeps latency low and cost predictable, and it is good enough for most summarising, extraction and classification work.
Move to a larger or reasoning model only when you can point to a specific task where the fast model is not good enough, such as complex multi-step reasoning or subtle writing. Reserve the strongest model for that task instead of routing everything through it.
See the Model Guides on the blog for a deeper comparison of when each family makes sense.