Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #475

Temperature (Sampling)

Temperature is a sampling hyperparameter that controls how "confident" or "random" an LLM's next-token choices are — a single number with an outsized effect on generated text's character.

Formula

\[ P(w_i) = \frac{e^{z_i/T}}{\sum_j e^{z_j/T}} \]

This is exactly the softmax formula from Softmax Function, with every logit \(z_i\) divided by a temperature \(T\) before exponentiating. \(T=1\) reproduces standard softmax exactly.

What Different Temperature Values Do

TemperatureEffect on the DistributionGenerated Text Character
\(T \to 0\)Sharpens toward the single highest-probability token (approaches argmax)Highly deterministic, repetitive, "safe"
\(T = 1\)Unchanged — the model's raw learned distributionBalanced
\(T > 1\)Flattens the distribution, making less-likely tokens relatively more probableMore random, varied, sometimes less coherent

Numerical Example

Logits \([2.0, 1.0, 0.1]\) at \(T=1\) (standard softmax, from Softmax Function): \([0.659, 0.242, 0.099]\). At \(T=0.5\) (dividing logits by 0.5, i.e. doubling them to \([4.0,2.0,0.2]\) before softmax): the distribution sharpens considerably, e.g. approximately \([0.843, 0.140, 0.017]\) — the top token becomes much more dominant. At \(T=2\) (halving the logits to \([1.0,0.5,0.05]\)): the distribution flattens toward more uniform, e.g. approximately \([0.475, 0.288, 0.237]\).

Code

import torch
import torch.nn.functional as F

logits = torch.tensor([2.0, 1.0, 0.1])

for T in [0.5, 1.0, 2.0]:
    scaled_probs = F.softmax(logits / T, dim=0)
    print(f"T={T}: {scaled_probs}")
# T=0.5: sharper, more confident distribution
# T=1.0: the model's original, unmodified distribution
# T=2.0: flatter, more uniform, more "random" distribution

When to Use Which Temperature

Task TypeTypical Temperature
Factual Q&A, code generation, mathLow (0–0.3) — favors reliability and determinism
Creative writing, brainstormingHigher (0.7–1.2) — favors variety and novelty

Common Mistakes

  • Setting temperature very high expecting purely "more creative" output without downside — excessive temperature can degrade coherence and factual reliability, not just increase variety.
  • Confusing temperature with top-k or top-p sampling — temperature reshapes the entire probability distribution's sharpness; top-k/top-p (next two notes) instead restrict which tokens are even eligible to be sampled from, a distinct and complementary mechanism.

Interview Relevance

Q: "Why would you use a low temperature for a code-generation task but a higher one for creative writing?" Code generation typically benefits from reliability and correctness — a low temperature sharpens the distribution toward the model's most confident (and typically most likely to be syntactically/logically correct) predictions. Creative writing benefits from variety and novelty — a higher temperature flattens the distribution, giving less-likely but potentially more interesting or original word choices a real chance of being sampled.

Practice Question

As \(T \to 0\), what does the sampling process approach — a random draw, or a deterministic choice? Explain using the softmax formula.

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →