Foundational Deep Learning interview questions — the ones almost every interview loop opens with. Each answer is explained in full, not just stated, so you can adapt it to however the question is actually phrased.
Q1. What's the actual difference between AI, Machine Learning, and Deep Learning?
AI is the broadest goal — building systems that perform tasks requiring intelligence. Machine Learning is a specific approach to AI: instead of hand-coding rules, a system learns patterns from data. Deep Learning is a specific approach to ML: it uses multi-layered neural networks to automatically learn hierarchical feature representations directly from raw data, rather than relying on hand-engineered features the way many classical ML algorithms do. So the relationship is nested: DL ⊂ ML ⊂ AI, each one a more specific technique within the broader one before it.
Q2. Why do we need activation functions? What would happen without them?
Without a non-linear activation function, stacking any number of linear layers collapses mathematically into a single equivalent linear transformation — no matter how many layers you add, the network could only ever represent linear relationships between input and output. Activation functions (ReLU, sigmoid, tanh, and others) introduce non-linearity between layers, which is what actually gives a deep network its expressive power to approximate complex, non-linear functions.
Q3. Explain forward propagation and backpropagation in your own words.
Forward propagation is the process of passing input data through the network layer by layer, applying each layer's weights and activation function, to produce a final output/prediction. Backpropagation is what happens next: using the chain rule, the network computes how much each individual weight contributed to the final loss (error), working backward from the output layer to the input layer — this gives the gradient for every weight, which the optimizer then uses to update the weights and reduce the loss.
Q4. What is the vanishing gradient problem, and why does it happen?
During backpropagation, gradients are computed by multiplying many partial derivatives together via the chain rule, layer by layer. If those individual derivatives are consistently small (e.g. sigmoid's derivative maxes out at 0.25), the product shrinks exponentially as it's propagated back through many layers — by the time the gradient reaches early layers, it can be so close to zero that those layers barely update at all, effectively stalling learning. This is one of the key reasons ReLU-family activations (which don't saturate for positive inputs) and techniques like residual connections became standard for deep networks.
Q5. What's the difference between a parameter and a hyperparameter?
A parameter is a value the model learns automatically during training — weights and biases. A hyperparameter is a value set by the person building the model before training begins, and it's not learned from data — learning rate, batch size, number of layers, and dropout rate are all hyperparameters. Choosing good hyperparameters is itself a search problem, covered in depth in the Hyperparameter Tuning category.
Q6. Why do we typically use mini-batch gradient descent instead of full-batch or pure single-example SGD?
Full-batch gradient descent (using the entire dataset per update) gives an accurate gradient estimate but is computationally expensive and memory-intensive per step, especially for large datasets. Pure single-example SGD is cheap per step but the gradient estimate is very noisy, making convergence erratic. Mini-batch gradient descent is the practical middle ground — it gets a reasonably accurate, much less noisy gradient estimate than single-example SGD, while still being computationally efficient and able to leverage parallel hardware (GPUs) effectively, which is exactly why it's the near-universal default in practice.
Q7. What is overfitting? Name three ways to reduce it.
Overfitting is when a model learns patterns specific to the training data — including its noise and idiosyncrasies — rather than the genuine underlying relationship, resulting in strong training performance but poor generalization to new, unseen data. Common mitigations: (1) regularization (L1/L2 weight penalties, dropout), (2) data augmentation to effectively increase training data diversity, and (3) early stopping — halting training once validation performance stops improving, before the model starts memorizing training-specific noise.
Q8. Explain the bias-variance tradeoff.
Bias is error from a model being too simple to capture the true underlying pattern (systematic underfitting). Variance is error from a model being too sensitive to the specific training data it happened to see, fitting noise rather than signal (overfitting). These two sources of error typically move in opposite directions as model complexity changes — increasing complexity usually reduces bias but increases variance, and vice versa. The goal is finding the complexity level that minimizes their combined effect on generalization error, not driving either one to zero in isolation.
Q9. What's the difference between an epoch, a batch, and an iteration?
An epoch is one complete pass through the entire training dataset. A batch is a subset of the training data processed together in one forward/backward pass. An iteration is one single update step — processing exactly one batch and updating the weights once. If a dataset has 10,000 examples and batch size is 100, one epoch consists of 100 iterations.
Q10. Why do we normalize or standardize input data before training?
Features on very different scales (e.g. one feature ranging 0-1, another ranging 0-100,000) can cause the loss surface to become elongated and poorly conditioned, making gradient descent converge slowly or unstably, since the same learning rate affects each feature's weight very differently. Normalizing/standardizing puts all features on a comparable scale, producing a better-conditioned loss surface that gradient descent navigates far more efficiently and reliably.