Coding Hubs School of AI – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 8448811540
Back to Deep Learning Notes
Topic #543

PyTorch Loss Functions

A practical reference catalog of PyTorch's built-in loss functions — every one already covered conceptually in the Loss Functions category, with exact syntax and the input-format details that most often trip people up.

Common Loss Functions, Quick Reference

LossPyTorch ClassExpectsConcept Note
MSEnn.MSELoss()Raw values for both prediction and targetMean Squared Error
MAEnn.L1Loss()Raw values for bothMean Absolute Error
Binary cross-entropynn.BCEWithLogitsLoss()Raw logits (NOT sigmoid-activated) + float labelsBinary Cross-Entropy
Categorical cross-entropynn.CrossEntropyLoss()Raw logits (NOT softmax-activated) + integer class indicesCategorical Cross-Entropy
Huber lossnn.HuberLoss(delta=1.0)Raw values for bothHuber Loss
KL divergencenn.KLDivLoss()Log-probabilities (first arg) + probabilities (second arg)KL Divergence Loss

Code — The Two Most Error-Prone Losses

import torch
import torch.nn as nn

# CrossEntropyLoss: raw logits + INTEGER class indices (not one-hot, not softmax-applied)
logits = torch.tensor([[2.0, 0.5, -1.0]])
labels = torch.tensor([0])                # integer index, not [1, 0, 0]
loss = nn.CrossEntropyLoss()(logits, labels)

# BCEWithLogitsLoss: raw logits + FLOAT labels (0.0 or 1.0)
logits_binary = torch.tensor([1.5])
labels_binary = torch.tensor([1.0])        # float, not int
loss_binary = nn.BCEWithLogitsLoss()(logits_binary, labels_binary)

Common Mistakes

  • Applying softmax/sigmoid manually before passing predictions to CrossEntropyLoss/BCEWithLogitsLoss — as flagged throughout this hub, this double-applies the activation and corrupts gradients.
  • Passing float labels to CrossEntropyLoss (which needs integer class indices) or integer labels to BCEWithLogitsLoss (which needs floats) — PyTorch will often raise a clear type error here, but it's a common first-time mistake.
  • Using reduction='sum' when 'mean' (the default) was intended, or vice versa — this changes the loss's effective scale relative to the learning rate, silently affecting training dynamics.

Interview Relevance

Q: "Why does nn.CrossEntropyLoss expect raw logits rather than softmax probabilities as input?" It applies log_softmax internally, computed in a numerically stable, combined way (avoiding the precision issues of computing softmax and then taking its log as two separate steps). Passing already-softmaxed probabilities would apply softmax twice, producing a mathematically incorrect loss and corrupted gradients.

Practice Question

For a binary classification model's raw output logit of 2.3 and a true label of 1, which PyTorch loss class would you use directly, without any manual activation applied first?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →