Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +918448811320
Back to Deep Learning Notes
Topic #270

Mean Absolute Error

Mean Absolute Error (MAE) measures the average magnitude of prediction errors, without regard to direction — it's one of the two default regression losses, and its defining trait is that it treats every unit of error equally, unlike MSE.

Formula

\[ \text{MAE} = \frac{1}{n}\sum_{i=1}^n |y_i - \hat{y}_i| \]

Numerical Example

True values \([10, 20, 30]\), predictions \([12, 18, 35]\). Errors: \(|10-12|=2\), \(|20-18|=2\), \(|30-35|=5\).

\[ \text{MAE} = \frac{2+2+5}{3} = 3.0 \]

Why "Absolute," Not "Squared"

Taking the absolute value (rather than squaring) means a large error contributes proportionally, not quadratically, to the total loss. This makes MAE substantially more robust to outliers than MSE (next note) — one wildly wrong prediction doesn't dominate the loss the way it would if errors were squared.

The Gradient Subtlety

MAE's derivative is a constant \(\pm1\) (the sign of the error) everywhere except exactly at zero error, where it's undefined (a sharp corner, like ReLU's kink). This means gradient descent on MAE takes a step of the same size regardless of how large the error is — unlike MSE, whose gradient scales with the error itself. In practice this can make MAE-trained models converge less smoothly very close to the optimum, which partly motivates Huber Loss (covered later in this category) as a middle ground.

Code

import numpy as np
import torch
import torch.nn as nn

y_true = np.array([10, 20, 30])
y_pred = np.array([12, 18, 35])
mae = np.mean(np.abs(y_true - y_pred))
print(mae)   # 3.0

loss_fn = nn.L1Loss()   # PyTorch's name for MAE loss
y_true_t = torch.tensor([10.0, 20.0, 30.0])
y_pred_t = torch.tensor([12.0, 18.0, 35.0])
print(loss_fn(y_pred_t, y_true_t))   # tensor(3.)

Common Mistakes

  • Defaulting to MSE for every regression task without considering whether outliers should be down-weighted — MAE is often the better choice when your data has occasional large, less-trustworthy errors (e.g. sensor glitches).
  • Forgetting MAE's gradient doesn't shrink as predictions get closer to correct — this can cause a small amount of oscillation right at convergence, which some optimizers handle better than others.

Interview Relevance

Q: "When would you choose MAE over MSE for a regression loss?" When the dataset likely contains outliers or noisy extreme values you don't want to dominate training — MAE's linear (not quadratic) penalty means a single large error contributes proportionally to the loss, not disproportionately, unlike MSE.

Practice Question

For true values \([5, 10]\) and predictions \([5, 20]\), compute MAE. Compare this to what MSE would give for the same errors (compute it too) — which loss is more affected by the single large error?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →