Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #271

Mean Squared Error

Mean Squared Error (MSE) is the most common regression loss — and as shown in Maximum Likelihood Estimation, it isn't an arbitrary choice: minimizing it is mathematically equivalent to maximum likelihood estimation under a Gaussian noise assumption.

Formula

\[ \text{MSE} = \frac{1}{n}\sum_{i=1}^n (y_i-\hat{y}_i)^2 \]

Numerical Example

Using the same values as Mean Absolute Error: true \([10,20,30]\), predicted \([12,18,35]\). Squared errors: \((-2)^2=4\), \(2^2=4\), \((-5)^2=25\).

\[ \text{MSE} = \frac{4+4+25}{3} \approx 11.0 \]

Compare this to MAE's 3.0 for the exact same predictions — the one large error (5) contributes 25 out of the total 33, over 75% of the loss, while it was only \(5/9 \approx 56\%\) of MAE's total. This is the direct numerical illustration of MSE's greater sensitivity to outliers.

Why Squaring Matters — The Gradient

\[ \frac{\partial}{\partial \hat{y}_i}(y_i-\hat{y}_i)^2 = -2(y_i-\hat{y}_i) \]

Unlike MAE's constant-magnitude gradient, MSE's gradient is proportional to the error itself — a large error produces a large gradient (pushing weights to correct it quickly), and a small error produces a small, gentle gradient (avoiding overshoot near the optimum). This smooth, error-proportional gradient is why MSE is usually easier to optimize with standard gradient descent than MAE.

Code

import numpy as np
import torch.nn as nn
import torch

y_true = np.array([10, 20, 30])
y_pred = np.array([12, 18, 35])
mse = np.mean((y_true - y_pred) ** 2)
print(mse)   # 11.0

loss_fn = nn.MSELoss()
y_true_t = torch.tensor([10.0, 20.0, 30.0])
y_pred_t = torch.tensor([12.0, 18.0, 35.0])
print(loss_fn(y_pred_t, y_true_t))   # tensor(11.)

Common Mistakes

  • Reporting raw MSE as a business-facing metric — its units are squared (e.g. "squared dollars"), which is rarely intuitive; RMSE (next note) restores interpretable units.
  • Using MSE on data known to have significant outliers without considering the alternatives — it will let a handful of extreme errors dominate training.

Interview Relevance

Q: "Why is MSE the 'natural' loss for regression, rather than an arbitrary convenient choice?" Assuming prediction errors are normally (Gaussian) distributed, minimizing MSE is mathematically equivalent to maximum likelihood estimation — the specific loss formula falls directly out of that probabilistic assumption, as derived in Maximum Likelihood Estimation, rather than being chosen arbitrarily.

Practice Question

A model has two candidate prediction sets with the same MAE but different MSE. What does the higher-MSE set likely indicate about the distribution of its individual errors?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →