
Temperature and Sampling, Explained
Temperature and top-p control how “random” a model’s output is. For most apps the defaults are fine, but knowing the dials helps when output feels too wild or too robotic.

Temperature in one line 🌡️
Lower temperature (closer to 0) makes answers predictable and focused; higher temperature (up to 1 or 2) makes them varied and creative.
| Task | Temperature | Why |
|---|---|---|
| Classification / extraction | 0 – 0.3 | Consistency matters |
| General chat | 0.7 | Balanced |
| Brainstorming / creative | 0.9 – 1.2 | More variety |
And top-p? 🎯
Top-p is an alternative sampler that only considers the most likely tokens that together cover probability p. As a rule of thumb, change one dial, not both.
Practical advice 💡
Start at the default. If answers are repetitive for structured tasks, lower the temperature. If they feel flat for creative work, raise it. Change one parameter at a time so you know what did it.
Frequently asked questions ❓
What is the default temperature?
0.7 for most chat-style calls, which suits general conversation well.
Does temperature affect cost?
No, it changes the style of output, not the number of tokens.
Should I use temperature or top-p?
Use temperature; leave top-p at its default unless you have a specific reason not to.
Comments (0)