
Handle Timeouts Gracefully
Reasoning models and peak hours both make latency unpredictable. The difference between a flaky integration and a robust one is how you handle a slow response.

Set a real timeout ⏲️
Most SDKs default to long timeouts. Set one that matches your model: a fast model should answer in a few seconds, while reasoning models may need 30 seconds or more.
| Model type | Suggested timeout | Notes |
|---|---|---|
| Fast / flash | 10–15s | Usually answers in 1–3s |
| Standard | 30s | Leave headroom for queues |
| Reasoning | 60s+ | Longer thinking time |
Retry with backoff 🔁
- Retry only idempotent requests or re-send the full prompt
- Start with a short delay (500ms) and double it
- Cap the number of retries at three
- Give up gracefully with a friendly error
Tell the user what’s happening 💬
For user-facing apps, show “still thinking…” rather than a spinner that never resolves. If a retry finally fails, respond with something useful instead of a bare error.
Frequently asked questions ❓
How many retries should I use?
Two or three with exponential backoff. Beyond that you are usually just extending the outage.
Should all errors be retried?
No. Only retry timeouts and transient errors. Do not retry authentication or invalid-request failures.
Does a timeout waste credits?
No. You only pay for requests the model actually processed.
Comments (0)