
Log Every API Call Without Slowing Down
When something goes wrong in production, the first question is always “what did the model actually receive?” A thin logging layer answers that instantly.

Wrap, don’t fork 🪢
Instead of sprinkling log statements through your code, wrap the SDK client in a small interceptor that records the essentials: model, input size, latency and status.
What to record 📝
| Field | Why it matters | Privacy tip |
|---|---|---|
| model + tokens | Cost and behaviour | Fine to store |
| latency | Spot slow models | Fine to store |
| prompt | Debug bad answers | Hash or redact PII |
| error | See failures fast | Store stack traces |
Keep it fast ⚡
Write logs asynchronously — push to a queue and flush in the background. Never block the API response on a disk write. If your framework has middleware, use it.
Sample in production 🎛️
Logging 100% of traffic is fine at low volume, but once you scale, sample to 10% and keep full logs only for errors. You keep the signal and drop the noise.
Frequently asked questions ❓
Where should I store the logs?
Anywhere you already log — stdout, a logging service or a database. Structured JSON lines are easiest to search.
Is it OK to log full prompts?
Only if you control the data and it contains no customer secrets. Otherwise redact or hash sensitive fields.
Does logging change latency?
Done correctly, the overhead is a few microseconds per call because the write happens in the background.
Comments (0)