Reconstruction loss measures how well a model can rebuild its original input after compressing and decompressing it — the core training signal for autoencoders, and one half of every VAE's loss (alongside KL divergence from the previous note).
The Setup
An autoencoder passes an input \(\mathbf{x}\) through an encoder (compressing it into a smaller latent representation \(\mathbf{z}\)) and then a decoder (expanding \(\mathbf{z}\) back into a reconstruction \(\hat{\mathbf{x}}\)). Reconstruction loss compares \(\hat{\mathbf{x}}\) directly to the original \(\mathbf{x}\) — there's no external label at all; the input is the target, which is exactly why autoencoders are a form of self-supervised/unsupervised learning (see Types of Learning).
Two Common Formulas, Chosen by Data Type
| Data Type | Typical Reconstruction Loss | Why |
|---|---|---|
| Continuous, unbounded (e.g. raw sensor readings) | MSE | Matches a Gaussian noise assumption, same reasoning as Mean Squared Error |
| Pixel values normalized to [0,1] (e.g. MNIST-style images) | Binary cross-entropy | Treats each pixel as a Bernoulli probability — matches the MLE derivation from Maximum Likelihood Estimation under a per-pixel Bernoulli assumption |
Numerical Example
A tiny autoencoder reconstructing a 3-pixel input \(\mathbf{x}=[0.9, 0.1, 0.5]\) as \(\hat{\mathbf{x}}=[0.8, 0.2, 0.5]\) using MSE: \(L = (0.9-0.8)^2+(0.1-0.2)^2+(0.5-0.5)^2 = 0.01+0.01+0 = 0.02\) (summed) or \(0.00\overline{6}\) (averaged over 3 pixels).
Code
import torch
import torch.nn as nn
class TinyAutoencoder(nn.Module):
def __init__(self):
super().__init__()
self.encoder = nn.Linear(784, 32) # compress
self.decoder = nn.Linear(32, 784) # reconstruct
def forward(self, x):
z = torch.relu(self.encoder(x))
x_hat = torch.sigmoid(self.decoder(z)) # sigmoid, since pixels are in [0,1]
return x_hat
model = TinyAutoencoder()
x = torch.rand(1, 784) # a flattened, normalized image
x_hat = model(x)
recon_loss = nn.BCELoss()(x_hat, x) # BCE, since pixels are treated as [0,1] probabilities
print(recon_loss)
Reconstruction Loss's Role in a VAE
These two terms pull in different directions: reconstruction loss alone would let the encoder use the latent space however it likes, as long as decoding works well; the KL term (from KL Divergence Loss) constrains that latent space to stay close to a well-structured prior. Balancing both is exactly what the VAE architecture, covered fully in the Generative Deep Learning category, is built around.
Common Mistakes
- Using MSE reconstruction loss on pixel data normalized to [0,1] by default, without considering BCE — for this specific data range, BCE's probabilistic framing is often the more principled and empirically better-performing match.
- Forgetting reconstruction loss requires no external labels at all — this is precisely what makes autoencoders self-supervised, and it's an easy point to lose track of when reasoning about what "training data" an autoencoder actually needs.
Interview Relevance
Q: "What's the training signal for an autoencoder, if there are no labels?" The input itself serves as the target — reconstruction loss compares the decoder's output directly against the original input the encoder compressed. This makes autoencoder training a form of self-supervised learning: the "label" is generated automatically from the data, with no human annotation required.
Key Takeaways — Loss Functions
- MAE, MSE and RMSE cover regression, with MAE more outlier-robust and MSE more optimization-friendly; Huber Loss deliberately combines both.
- Binary and categorical cross-entropy are the classification defaults, paired with sigmoid and softmax respectively — always double-check whether your framework's loss function expects raw logits or already-activated probabilities.
- Sparse categorical cross-entropy is not a different formula — only a different label format (integer index vs. one-hot).
- Focal Loss down-weights easy examples to fix severe class imbalance; contrastive and triplet losses shape embedding spaces directly rather than predicting fixed labels.
- KL divergence and reconstruction loss are the two building blocks behind every VAE, and reconstruction loss alone drives every plain autoencoder.
Next: Optimization & LR Scheduling covers exactly how these loss values get turned into weight updates — SGD through AdamW, and every learning-rate schedule used in practice.
Practice Question
You're building an autoencoder to compress and reconstruct raw audio waveform samples (continuous values, not bounded to [0,1]). Which reconstruction loss — MSE or BCE — is the more appropriate choice, and why?