Learn Deep Learning from first principles โ math, neural networks, CNNs, sequence models, Transformers, LLMs, PyTorch and production deployment.
AI vs ML vs DL, why deep learning works now, the DL workflow and its real limitations.
Scalars to eigenvectors โ the matrix and tensor operations every neural network runs on.
Derivatives, gradients and computational graphs โ the machinery behind learning.
Distributions, likelihood, entropy and KL divergence โ the language of loss functions.
From the biological neuron to the MLP โ layers, weights, forward propagation and the full training flow.
Every major activation function โ formula, graph, derivative and where it's actually used.
Regression, classification and advanced losses โ with formulas and worked numbers.
SGD to AdamW, and every learning-rate schedule that actually gets used in practice.
The chain rule applied to a real network, with a full numerical worked example.
Epochs, checkpoints, early stopping โ and the underfitting/overfitting tradeoff.
L1/L2, dropout, weight decay and augmentation compared side by side.
BatchNorm, LayerNorm, GroupNorm and RMSNorm โ formulas and when each is used.
Classification, regression and modern DL metrics โ from confusion matrix to BLEU and mAP.
Convolution, kernels, stride, padding and pooling โ with a manual convolution example.
LeNet to ConvNeXt โ the innovation, and the problem, behind every landmark CNN.
Detection and segmentation families โ R-CNN through YOLO, FCN through Mask R-CNN.
Sequential data, hidden state, unrolling, and why plain RNNs struggle with long sequences.
Gates, cell state and the full equations for LSTM and GRU, compared directly.
Encoder-decoder, the context-vector bottleneck, and attention built from first principles (Q/K/V).
The full architecture โ encoder, decoder, positional encoding and residual connections โ with data flow.
Tokenization to embeddings โ Word2Vec, GloVe, BERT, RoBERTa, T5 and GPT-style models.
Pretraining to RLHF/DPO, context windows, sampling and KV-cache โ how LLMs actually run.
Autoencoders, VAEs and GANs โ generative vs discriminative, explained from the ground up.
Forward noising, reverse denoising, latent diffusion and Stable Diffusion.
Pretrained models, freezing layers, and fine-tuning vs feature extraction.
LoRA, QLoRA, adapters and quantization โ with the LoRA math worked out.
Pretext tasks, contrastive learning (SimCLR, MoCo) and masked modeling.
Few-shot, meta-learning, distillation, federated learning and Mixture of Experts.
Tensors, autograd, nn.Module, DataLoader and full training loops โ hands-on.
Sequential and Functional APIs, callbacks, and how TensorFlow compares to PyTorch.
What to tune, and how โ grid search, random search, Bayesian optimization and Optuna.
The full lifecycle โ problem definition through monitoring โ applied end to end.
TorchScript, ONNX, FastAPI, Docker and cloud deployment for real inference traffic.
Experiment tracking, model registries, drift detection and distributed training.
Embeddings, vector databases, RAG, agents and long-context/reasoning models โ from a DL lens.
Reading papers, baselines, SOTA, ablations, reproducibility, FLOPs and benchmarking.
Topic-wise Deep Learning interview questions with fully explained answers.
Implementation-oriented exercises across networks, CNNs, RNNs, transformers and PyTorch.
End-to-end builds โ image classifiers through fine-tuned LLMs โ with deployment.
No notes found. Try a different search term, or browse all Deep Learning notes.
What Artificial Intelligence is, its sub-fields, and how deep learning fits inside it โ the map for this whole notes hub.
The exact containment relationship between AI, Machine Learning and Deep Learning, with a side-by-side comparison table.
Manual feature engineering vs automatic feature learning, compared side by side with a pipeline diagram.
From the 1958 Perceptron through two AI winters to the 2012 breakthrough and the modern LLM era.
The three pillars โ data, compute and algorithmic fixes โ that made deep networks finally trainable at scale.
Real-world deep learning applications across computer vision, NLP, healthcare, finance and generative AI.
Supervised, unsupervised, semi-supervised, self-supervised and reinforcement learning in a deep learning context.
The end-to-end deep learning project workflow, from problem definition to deployment and monitoring.
Data, model, loss function, optimizer and hardware โ the components every trained deep learning system depends on.
Data hunger, compute cost, overfitting, vanishing gradients and interpretability โ the real engineering challenges.
A practical decision guide for when deep learning's advantages outweigh its costs, and when they don't.
The three basic mathematical objects behind every neural network โ scalars, vectors and matrices โ with notation and shapes explained.
Tensors as the n-dimensional generalization of scalars, vectors and matrices, with rank, shape and PyTorch code.
Element-wise matrix addition, its formula, and its role in bias terms and ResNet residual connections.
The shape rule, a full worked example, and why matrix multiplication is the core operation of every linear layer.
The dot product formula, geometric intuition, and its role in cosine similarity and attention scores.
What matrix transpose does, its key properties, and its role in backpropagation and attention.
The matrix inverse, the 2x2 formula, and why deep learning uses iterative optimization instead of computing inverses.
The determinant formula, its geometric meaning as area/volume scaling, and its link to invertibility.
L1, L2 and L-infinity norms compared, with their role in regularization and gradient clipping.
What eigenvalues are, the characteristic equation, a worked numerical example, and their role in PCA.
What eigenvectors are, how to find them, and how they power Principal Component Analysis.
The formal definition of a vector space, basis, span and dimension, connected to feature and embedding spaces.
What linear transformations are, common examples, and why they motivate the need for activation functions.
The broadcasting rule in NumPy and PyTorch, worked examples, and common shape bugs to avoid.
What functions and limits are, and why calculus is needed at all to train a neural network.
The formal definition of a derivative, common rules, geometric meaning, and PyTorch autograd.
How partial derivatives isolate the sensitivity of a multi-variable loss to a single weight.
The chain rule formula, a worked example, and why it is the mathematical basis of backpropagation.
What the gradient vector is, why it points uphill, and how gradient descent uses it to minimize loss.
The shape and interpretation of a gradient vector for a real weight vector, with a PyTorch example.
The Jacobian matrix for multi-input, multi-output functions, and why every network layer needs one.
The Hessian matrix of second derivatives, what it reveals about minima, maxima and saddle points.
How computational graphs represent a function's operations, and how PyTorch uses them for autograd.
A complete end-to-end summary connecting derivatives, the chain rule and gradients to how a network actually trains.
Core probability definitions and axioms, and why a classifier's softmax output must satisfy them.
Discrete vs continuous random variables, and how dataset labels and weight initialization use them.
PMFs, PDFs, the Bernoulli, categorical and normal distributions, and their roles across deep learning.
The formula for expected value, and how training loss is literally an empirical expected value.
Variance and standard deviation formulas, and why they motivate feature standardization and batch normalization.
The covariance formula, the covariance matrix, and its role in PCA and dimensionality reduction.
The conditional probability formula, a spam-filter worked example, and independence.
Bayes' theorem formula, prior/likelihood/posterior, and a full worked medical-test example.
The difference between probability and likelihood, with a coin-flip worked example.
MLE explained, and a full derivation of cross-entropy and MSE loss from maximum likelihood principles.
The entropy formula, worked examples, and how it measures uncertainty in a probability distribution.
The cross-entropy formula, a full worked example, and PyTorch's CrossEntropyLoss explained.
The KL divergence formula, its relationship to cross-entropy and entropy, and its use in VAEs and RLHF.
The biological neuron structure that inspired artificial neural networks, mapped part by part.
The full artificial neuron computation โ weighted sum plus activation โ with a worked numerical example.
The first mathematical neuron model from 1943, and how it modeled logical AND and OR without learning.
The 1958 Perceptron model, its linear decision boundary, and what makes it different from earlier neuron models.
The Perceptron's weight update rule, a full worked example, and the Perceptron Convergence Theorem.
Why a single Perceptron cannot learn XOR, proven algebraically, and how a hidden layer fixes it.
The MLP architecture that solves XOR, with a diagram and full PyTorch implementation.
Input, hidden and output layers, depth vs width, and common architecture patterns.
What each layer type does, and how to correctly size and activate the output layer for any task.
What weights and bias represent, weight matrix shapes, and why zero-initialization fails.
The precise distinction between learned parameters and chosen hyperparameters, with clear examples.
The forward propagation formula, a full numerical walkthrough, and matching PyTorch code.
What a loss function measures, MSE vs cross-entropy introduced, and loss vs evaluation metric.
A conceptual introduction to backpropagation as efficient gradient computation via the chain rule.
The gradient descent update rule, why learning rate matters, and a complete PyTorch training step.
A complete worked example of one full training step, from forward pass to weight update, verified in PyTorch.
The original 1958 Perceptron activation function, and why its zero gradient makes it unusable for backpropagation.
The linear (identity) activation, why it collapses network depth, and where it's actually used in regression outputs.
The sigmoid formula, derivative, graph, and why it causes vanishing gradients in deep networks.
The tanh formula, derivative, and how it improves on sigmoid while sharing its saturation problem.
The ReLU formula, why it fixes vanishing gradients, and the dying ReLU problem explained.
How Leaky ReLU's small negative slope fixes the dying ReLU problem, with formula and code.
How PReLU makes the negative slope a learnable parameter instead of a fixed hyperparameter.
The ELU formula, its smooth exponential negative curve, and its tradeoffs versus ReLU.
The self-normalizing SELU activation, its strict requirements, and why it's used less often in practice.
The GELU formula, its probabilistic derivation, and why it's the default in Transformer architectures.
The Swish activation function, discovered via search, and its close relationship to GELU.
The softmax formula, a full worked numerical example, and the numerical stability trick every framework uses.
A direct comparison of sigmoid and tanh, and why both still appear inside LSTM and GRU gates.
A direct comparison of ReLU and Leaky ReLU, and when the dying ReLU problem actually justifies switching.
Why CNNs favor ReLU while Transformers favor GELU, compared side by side.
When to use sigmoid vs softmax for an output layer, based on whether classes are mutually exclusive.
The MAE formula, a worked example, and why it's more robust to outliers than MSE.
The MSE formula, a worked example, and its derivation from maximum likelihood under a Gaussian assumption.
The RMSE formula and why it's reported as a metric while MSE is used as the training loss.
The Huber loss formula, a worked example, and how it combines MSE and MAE's strengths.
The BCE formula, a worked example, and the critical BCELoss vs BCEWithLogitsLoss distinction.
The categorical cross-entropy formula, a worked example, and PyTorch's expected label format.
Why sparse categorical cross-entropy is mathematically identical to categorical cross-entropy, just a label format difference.
The Focal Loss formula and how it down-weights easy examples to fix severe class imbalance.
The contrastive loss formula, a worked example, and how it shapes embedding spaces for similarity.
The triplet loss formula, a worked example, and the hard negative mining challenge in practice.
How KL divergence is implemented as a practical loss function, with the PyTorch input-format gotcha explained.
MSE vs BCE reconstruction loss for autoencoders, and its role alongside KL divergence in a VAE.
The optimization problem formalized, convex vs non-convex loss surfaces, and iterative descent visualized.
How batch gradient descent uses the full dataset per update, and why it's rarely used in practice.
How SGD uses one random example per update, its noisy path, and why that noise can help.
The practical compromise between batch and stochastic gradient descent, and how it maps to PyTorch's DataLoader.
A deep dive on the learning rate hyperparameter, typical ranges, and its interaction with batch size.
The momentum formula, the physical rolling-ball analogy, and how it smooths oscillation in narrow ravines.
How Nesterov momentum's look-ahead gradient improves on standard momentum.
The AdaGrad formula, per-parameter adaptive learning rates, and its diminishing learning rate flaw.
How RMSProp fixes AdaGrad's diminishing learning rate problem with a decaying average.
The Adam optimizer's formula, bias correction explained, and why it became the default optimizer.
How AdamW fixes Adam's broken interaction with L2 regularization, with a full optimizer comparison table.
Why a fixed learning rate is a compromise, and an overview of the major scheduling strategies.
The step decay formula, a worked example, and its staircase-shaped learning rate curve.
The exponential decay formula and how it compares to step decay's sudden drops.
The cosine annealing formula, why its smooth shape is popular, and warm restarts explained.
Why gradually increasing the learning rate at the start of training stabilizes early updates.
The One Cycle policy's rise-fall-anneal shape and how it enables faster "super-convergence" training.
Why backpropagation exists โ the impractical cost of finite differences versus backprop's efficiency.
How the chain rule applies to a multi-layer network, with local gradients reused across layers.
What values must be cached during the forward pass for backpropagation to use later.
The recursive formula for computing each layer's error signal during the backward pass.
The outer product formula for computing weight and bias gradients from a layer's error signal.
How backpropagation's computed gradients are handed off to a separate, swappable optimizer.
A complete numerical backpropagation example through a 2-layer network, verified with PyTorch autograd.
The full mechanism behind vanishing gradients, a numerical demonstration, and every solution mapped out.
The mechanism behind exploding gradients, a numerical demonstration, and how to recognize it in practice.
The gradient clipping formula, a worked example, and why it's standard practice for RNNs and Transformers.
Why datasets are split into three parts, typical ratios, and the data leakage traps to avoid.
The precise definitions of epoch, batch and iteration, and how they relate mathematically.
The complete, annotated standard PyTorch training loop structure, piece by piece.
The complete validation loop structure, model.eval(), and why it differs from the training loop.
What to save in a training checkpoint, periodic vs best-model checkpointing, with full PyTorch code.
The patience-based early stopping algorithm, with full code and how to choose the patience value.
state_dict vs whole-model saving in PyTorch, and how to correctly load partial weights for transfer learning.
The signature of underfitting, its common causes, and the specific fixes for each.
The signature of overfitting, why it happens, and every major fix explained.
The formal bias-variance decomposition of model error, and how it explains underfitting and overfitting.
What generalization means, the generalization gap, and why deep networks generalize despite huge parameter counts.
The general idea behind regularization โ trading a little bias for a lot less variance.
The L1 regularization formula, why it produces sparse weights, and when it's the right choice.
The L2 regularization formula, why it shrinks weights smoothly, and why it's the common default.
The weight decay formula, its algebraic equivalence to L2 under SGD, and why it breaks under Adam.
How dropout randomly disables neurons during training, inverted dropout scaling, and why it works.
How data augmentation expands the effective training set, common techniques, and the label-preservation rule.
A complete side-by-side comparison of every regularization technique, and a practical decision guide.
The shared formula behind every normalization technique, and the shifting-activation-distribution problem they solve.
The BatchNorm formula, the critical training-vs-inference statistics split, and its small-batch weakness.
The LayerNorm formula, how it differs from BatchNorm, and why Transformers use it.
The InstanceNorm formula and why it suits style transfer by removing per-instance contrast.
The GroupNorm formula, how it generalizes InstanceNorm and LayerNorm, and its batch-size independence.
The RMSNorm formula, why it drops mean-centering, and a complete comparison of all 5 normalization techniques.
The 2x2 confusion matrix foundation โ TP, TN, FP, FN โ every classification metric is built from.
The accuracy formula and why it can be deeply misleading on imbalanced classification datasets.
The precision and recall formulas, their tradeoff, and how to choose which matters more by context.
The F1 score formula, why it uses a harmonic mean, and the F-beta generalization.
The specificity (true negative rate) formula and how it complements recall.
The ROC curve's TPR vs FPR plot across every threshold, with a worked code example.
The ROC-AUC metric, its probabilistic interpretation, and a worked numerical example.
The Precision-Recall curve and why it's more informative than ROC for imbalanced datasets.
The PR-AUC metric, its class-imbalance-dependent baseline, and when to prefer it over ROC-AUC.
The Rยฒ metric formula, a worked example, and how MAE/MSE/RMSE double as evaluation metrics.
The perplexity formula for language model evaluation, and its intuitive branching-factor interpretation.
The BLEU score formula, n-gram precision, the brevity penalty, and its known limitations.
The ROUGE score formula, why it emphasizes recall, and how it compares to BLEU.
The IoU formula, a worked bounding box example, and its role in object detection matching.
The Dice score formula, its exact relationship to IoU, and why it's popular in medical imaging.
How mAP combines IoU matching and per-class precision-recall into the standard object detection metric.
A high-level introduction to CNNs, the convolution/pooling/FC pipeline, and a minimal PyTorch example.
Why plain MLPs fail on images, and how sparse connectivity and parameter sharing fix it, with a parameter comparison.
How images are represented as tensors, pixel value ranges, and channel-first vs channel-last conventions.
The convolution formula, a worked numerical example, and its implementation in PyTorch.
The precise difference between a kernel and a filter, with classic edge-detection kernel examples.
What a feature map represents, and how multiple filters produce multiple stacked feature maps.
The stride parameter, its output size formula, and why larger strides act as built-in downsampling.
The padding parameter, valid vs same padding, and why it prevents edge-pixel underprocessing.
How multi-channel convolution combines input channels, and input vs output channel counts.
How receptive field grows across stacked layers, and why small kernels are stacked instead of large ones.
The general pooling concept, why it helps, and its zero-parameter downsampling benefit.
The max pooling formula, a worked numerical example, and why it suits feature detection.
The average pooling formula, a worked example, and when it's preferred over max pooling.
How GAP collapses entire feature maps, with a dramatic parameter-count comparison to flattening.
How flattening bridges spatial feature maps to fully connected layers, with a worked example.
The role of the fully connected layer at the end of a CNN, and why it often dominates parameter count.
A complete manual convolution example with padding and stride, verified against PyTorch's conv2d.
The original 1998 CNN for digit recognition, establishing the conv-pool-FC template.
The 2012 ImageNet breakthrough that kicked off the modern deep learning era.
The VGG architecture's uniform 3x3 kernel design philosophy, and its parameter count tradeoff.
How GoogLeNet's Inception modules and 1x1 convolutions achieved efficiency over VGG.
A deep dive into the Inception module's parallel multi-scale branches and concatenation.
How residual connections solved vanishing gradients, enabling networks over 100 layers deep.
How DenseNet's dense, concatenated connections maximize feature reuse across layers.
How depthwise separable convolutions make MobileNet efficient for mobile deployment.
How compound scaling balances depth, width and resolution together for efficient accuracy.
How Xception takes the Inception idea to its extreme with depthwise separable convolutions throughout.
How ConvNeXt modernized CNNs with Transformer-inspired design to match Vision Transformer performance.
The core image classification task, the standard CNN pipeline, and benchmark datasets.
Feature extraction vs fine-tuning for reusing pretrained CNNs, with PyTorch code for both.
The object detection task, one-stage vs two-stage detector families, and evaluation with mAP.
The general pixel-level segmentation task, and the three flavors: semantic, instance and panoptic.
Per-pixel class labeling without instance distinction, with a worked numerical example.
How instance segmentation distinguishes individual objects, built via detection plus masking.
How panoptic segmentation unifies "things" and "stuff" into one complete scene output.
Embedding-based face verification and identification using triplet and contrastive loss.
The two-stage OCR pipeline of text detection and sequence-based text recognition.
The CNN encoder plus RNN/Transformer decoder pipeline for generating image descriptions.
Heatmap-based keypoint detection for reconstructing body pose from an image.
The original region-proposal-based object detector, and why it was extremely slow.
How Fast R-CNN shares CNN computation across regions using ROI Pooling.
How the Region Proposal Network made object detection fully end-to-end trainable.
How SSD predicts detections in a single pass using multi-scale feature maps.
The grid-based single-pass detection approach behind YOLO's real-time speed.
How FCN replaced fully connected layers with convolutions for practical semantic segmentation.
How Mask R-CNN adds instance segmentation to Faster R-CNN with ROI Align and a mask branch.
What makes data sequential, common examples, and why standard feedforward networks struggle with it.
The three requirements a sequence model needs, and how the RNN architecture satisfies all three.
The three weight matrices of an RNN cell, and how RNN architecture compares to a standard feedforward layer.
What an RNN's hidden state represents, its fixed-size bottleneck, and the batch-size tradeoff.
The recurrent connection formalized, and why the same weights must be shared across every time step.
How unrolling an RNN across time steps reveals its connection to backpropagation and network depth.
The complete RNN forward pass formulas, with a full 3-step numerical example verified in PyTorch.
How BPTT sums gradient contributions across time steps for RNN's shared weights, and truncated BPTT.
Why vanishing gradients hit RNNs especially hard over long sequences, with a numerical example.
Why RNNs are especially prone to exploding gradients, and why gradient clipping is nearly mandatory.
The complete list of RNN limitations, why LSTM/GRU exist, and the parallelization problem they still don't solve.
The core motivation for LSTM โ a cell state pathway that protects gradients across long sequences.
The five components of an LSTM cell and how they connect, with a diagram and PyTorch weight shapes.
The cell state update formula and why its largely-additive structure protects gradients across time.
How the LSTM hidden state is derived from the cell state, and why the two states are kept separate.
The LSTM forget gate formula, a worked example, and how it decides what memory to discard.
The LSTM input gate formula, a worked example, and how it decides how much new information to add.
The LSTM candidate state formula, why it uses tanh instead of sigmoid, and a worked example.
The LSTM output gate formula and how it controls what the cell state exposes as the hidden state.
All six LSTM equations assembled into one complete reference, with a summary table and diagram.
A complete numerical LSTM forward pass through one time step, verified against PyTorch.
LSTM's genuine strengths and remaining weaknesses, including the parallelization problem it never solved.
Real-world LSTM applications โ translation, speech recognition, forecasting โ with PyTorch code examples.
GRU's simplified structure compared to LSTM โ one state, two gates instead of two states, three gates.
The GRU update gate formula, how it merges LSTM's forget and input gates, with a worked example.
The GRU reset gate formula and how it shapes the candidate hidden state computation.
All four GRU equations assembled together, with a complete numerical worked example.
A complete side-by-side comparison of GRU and LSTM, with practical guidance on which to choose.
How splitting a model into an encoder and decoder decouples input and output sequence lengths.
The fixed-size context vector, why it's a bottleneck, and its role in basic Seq2Seq models.
The complete assembled Seq2Seq pipeline and autoregressive decoding, with a full code example.
How teacher forcing speeds up Seq2Seq training, and the exposure bias it introduces.
The complete list of basic Seq2Seq limitations that directly motivated the attention mechanism.
The core motivation for attention, and how it directly fixes the Seq2Seq context-vector bottleneck.
The Q/K/V framework explained via a search-engine analogy, with formulas and a numerical example.
How the dot product between query and key produces a raw relevance score, with a worked example.
The complete dot-product attention formula, from scores to softmax to weighted value sum.
Why attention scores are divided by the square root of dk, and how this keeps softmax gradients healthy.
How self-attention relates every token in a sequence directly, without an RNN's distance penalty.
How cross-attention lets a decoder query an encoder's outputs, formally solving the Seq2Seq bottleneck.
How running multiple attention heads in parallel lets a model capture different relationship types.
Why "Attention Is All You Need" removed recurrence entirely, and why parallelization matters so much.
The complete Transformer architecture at a glance, plus encoder-only, decoder-only and encoder-decoder variants.
One Transformer encoder layer's structure โ self-attention and feed-forward, each with residual and norm.
One Transformer decoder layer's structure โ masked self-attention, cross-attention, and feed-forward.
The sinusoidal positional encoding formula, why sine and cosine, with a worked example.
The position-wise feed-forward sublayer's formula and why it adds non-linear capacity to attention output.
How residual connections around every Transformer sublayer enable training very deep stacks.
Post-norm vs pre-norm placement in Transformer blocks, and why layer norm is used over batch norm.
How causal masking prevents the decoder from seeing future tokens, with the exact masking mechanism.
The complete Transformer data flow from input tokens to predicted output, with a full PyTorch implementation.
Classical text preprocessing steps and why modern Transformers rely on much less of them.
Word, character and subword tokenization compared, and why BPE became the modern standard.
The vocabulary size tradeoff and how it directly sizes a model's embedding and output layers.
Why one-hot word encoding fails โ huge dimensionality and zero notion of similarity.
How dense word embeddings capture semantic similarity, with the famous king-queen analogy.
The self-supervised idea behind Word2Vec, and a preview of its CBOW and Skip-Gram architectures.
The Continuous Bag of Words formula and how it predicts a target word from averaged context.
The Skip-Gram formula, and a direct comparison with CBOW including rare-word performance.
How GloVe learns embeddings from global co-occurrence statistics, compared to Word2Vec.
Why static embeddings can't distinguish word meanings, and how contextual embeddings fix this.
BERT's bidirectional architecture, Masked Language Modeling, and Next Sentence Prediction explained.
How RoBERTa improved on BERT using the identical architecture, purely through a better training recipe.
How T5 reframes every NLP task as text-to-text using a full encoder-decoder Transformer.
GPT's decoder-only architecture, next-token prediction, and a comparison against BERT.
What makes a model "large," emergent capabilities, and why LLMs aren't automatically chatbots.
The modern refinements (RMSNorm, RoPE, grouped-query attention) on top of the base Transformer decoder.
Why token counts matter for cost, context budget and latency, with tiktoken code.
The LLM input embedding layer at scale, and the weight-tying trick between input and output layers.
Learned positional embeddings vs Rotary Positional Embeddings (RoPE), and why RoPE became dominant.
The transformer block as the fundamental unit of LLM scale, and depth vs width tradeoffs.
The massive self-supervised pretraining stage, and a brief look at scaling laws.
The full formula for next-token prediction and the parallel-supervision trick causal masking enables.
How SFT reshapes a pretrained model using a much smaller, curated dataset of example responses.
How training on diverse instruction-phrased tasks teaches models to generalize to unseen instructions.
Why capability and alignment are distinct concerns, and what alignment actually aims to achieve.
The full 3-stage RLHF process โ reward model training and RL fine-tuning with a KL penalty.
How Direct Preference Optimization achieves RLHF-like results without a reward model or RL.
Why attention's quadratic cost limits context window size, with a practical token-budget example.
The temperature sampling formula and how it sharpens or flattens a model's output distribution.
The top-K sampling algorithm, a worked example, and its fixed-count limitation.
Nucleus (top-P) sampling, how it adapts to distribution shape, compared directly to top-K.
Why the KV cache is valid, and how it turns quadratic generation cost into linear cost.
The perplexity formula, a worked example, and what perplexity does and doesn't measure.
The core distinction between learning P(y|x) and learning the data distribution P(x) itself.
The encoder-bottleneck-decoder structure, why the bottleneck matters, and practical uses.
How training on corrupted inputs with clean targets produces more robust learned features.
How a sparsity penalty forces specialized, interpretable features even with a large latent space.
The VAE's probabilistic latent space, the reparameterization trick, and the full loss formula.
The GAN minimax formula, adversarial training loop, and why training is notoriously unstable.
The architectural guidelines that stabilized convolutional GAN training.
How conditioning both generator and discriminator on a label enables controlled generation.
StyleGAN's mapping network and multi-resolution style injection for fine-grained generation control.
Why diffusion models offer more stable training than GANs, at the cost of generation speed.
The fixed forward noising process formula, and the closed-form shortcut to any noise step.
The learned reverse denoising process and the full generation loop from pure noise to an image.
The simplified noise-prediction training objective diffusion models actually use in practice.
Why U-Net's skip connections and timestep conditioning suit it for diffusion noise prediction.
How cross-attention and classifier-free guidance steer diffusion generation toward a text prompt.
How running diffusion in a compressed VAE latent space makes generation computationally practical.
The complete Stable Diffusion pipeline assembling text encoding, VAE and U-Net conditioning.
The practical parameters shaping text-to-image output, prompt engineering, and reproducible generation.
A complete side-by-side comparison of diffusion models and GANs, and when to choose each.
Why transfer learning works, based on the general-to-specific pattern in learned representations.
Where pretrained models come from, why they save so much compute and data, and how to choose one.
The frozen-backbone feature extraction approach, when it works best, and full PyTorch code.
The requires_grad mechanic behind freezing, and progressive unfreezing as a training strategy.
How fine-tuning differs from feature extraction, and why it needs a much smaller learning rate.
A decision framework for choosing feature extraction, partial or full fine-tuning based on your data.
Handling systematic distribution shift between source and target domains beyond standard fine-tuning.
The real memory cost of full fine-tuning at LLM scale, and why it becomes prohibitive.
The core idea behind Parameter-Efficient Fine-Tuning and why training under 1% of parameters can work.
The LoRA formula, exact parameter savings math, and a full PyTorch implementation.
How QLoRA combines LoRA with 4-bit quantization to fine-tune huge models on one GPU.
The adapter bottleneck module formula, and a direct comparison to LoRA including inference latency.
How trainable virtual tokens at every attention layer steer a frozen model's behavior.
The simplest PEFT method, and a direct comparison against prefix tuning.
The quantization formula, a worked numerical example, and the precision/efficiency tradeoff.
NF4 and double quantization, the techniques that keep 4-bit quantization accurate.
The 8-bit quantization middle ground, and 8-bit optimizer states for training memory savings.
The precise three-way distinction between supervised, unsupervised and self-supervised learning.
What a pretext task is, common examples across vision and NLP, and why solving it well matters.
The InfoNCE loss formula and the core positive/negative pair recipe behind contrastive learning.
The SimCLR pipeline, its projection head trick, and why it needs large batch sizes.
How MoCo's momentum encoder and negative queue decouple contrastive learning from batch size.
MLM as a pretext task, and the 80/10/10 masking scheme that reduces train/fine-tune mismatch.
Why masked image modeling needs a much higher masking ratio than text, with MAE-style code.
The overarching goal of self-supervised learning, and linear probing as the standard evaluation.
In-context learning vs prototypical networks, and how LLMs achieve few-shot learning with zero updates.
How CLIP achieves zero-shot image classification, with full code and why it isn't magic.
The MAML algorithm's inner and outer loops, with formulas and a code sketch.
The distillation loss formula, why soft labels carry "dark knowledge," and full PyTorch code.
The catastrophic forgetting problem and the EWC regularization technique to combat it.
The FedAvg algorithm for training across decentralized, private data, with the aggregation formula.
Early, late and intermediate fusion strategies for combining multiple data modalities.
CLIP's contrastive training objective across images and text, with a full code example.
The MoE routing formula and how sparse activation decouples capacity from compute cost.
The three components of NAS and the search strategies used to automate architecture design.
Installing PyTorch with GPU support and verifying CUDA availability correctly.
Practical tensor creation, inspection, and the NumPy conversion memory-sharing gotcha.
Indexing, reshaping, view vs reshape, and the cat vs stack distinction, with code.
The classic silent shape-mismatch bug in real PyTorch code, with a worked debugging example.
The practical autograd API โ requires_grad, backward(), no_grad(), and gradient accumulation.
How PyTorch builds dynamic graphs on every forward pass, and retain_graph explained.
The base class for all PyTorch models, parameter registration, and train/eval modes.
A practical reference catalog of common PyTorch layer types with exact syntax.
Module vs functional activation forms, with a full quick-reference table.
A practical reference for PyTorch loss classes and their exact input format requirements.
A practical reference for PyTorch optimizers, including per-parameter-group learning rates.
The minimal Dataset interface โ __len__ and __getitem__ โ with a full working example.
DataLoader parameters, batching, shuffling, and custom collate functions for variable-length data.
The complete practical training loop with device management and memory-safe loss logging.
The complete practical validation loop with correctly weighted loss averaging.
Moving models and data to GPU correctly, device-mismatch errors, and mixed precision training.
Practical syntax for saving weights, full checkpoints, and best-model tracking.
Practical syntax for loading weights, checkpoints, and partial weights across devices.
Complete, runnable PyTorch code for feature extraction and partial fine-tuning.
A complete realistic custom Dataset loading images from disk with CSV labels.
When to write a custom training loop, with gradient accumulation and custom logging patterns.
Eager execution vs @tf.function graph compilation in modern TensorFlow.
The tf.constant vs tf.Variable distinction and GradientTape for automatic differentiation.
The three Keras model-building APIs previewed, and how Keras relates to TensorFlow.
Building linear layer stacks with the Sequential API, including a full CNN example.
Building branching architectures, residual connections and multi-input models with the Functional API.
Subclassing tf.keras.Model with call(), and the training argument for mode-dependent layers.
Built-in callbacks for early stopping, checkpointing and learning rate scheduling.
The compile() and fit() high-level training API, contrasted with PyTorch's manual loop.
The evaluate() and predict() methods, and the difference between them.
Keras's full-model-save default compared to PyTorch's state_dict-focused approach.
A complete side-by-side comparison of PyTorch and TensorFlow/Keras with practical decision guidance.
A complete catalog of every hyperparameter covered in this hub, and the general tuning philosophy.
The learning rate range test technique and how to diagnose too-high vs too-low learning rates.
The linear scaling rule, memory constraints, and the generalization tradeoff of batch size.
Why modern practice relies on early stopping rather than fixing epoch count manually.
A practical strategy for choosing network depth based on underfitting/overfitting symptoms.
Common width patterns like the funnel shape, and the width/data-size relationship.
Typical dropout rates, where to apply dropout, and diagnosing rate adjustments from symptoms.
Typical weight decay ranges and why AdamW is essential for meaningful weight decay tuning.
A practical decision guide for choosing between AdamW, SGD and RMSProp.
A practical decision guide for choosing activation functions by layer type and architecture.
The exhaustive grid search algorithm, its combinatorial explosion problem, and code.
Why random search often outperforms grid search for the same compute budget.
The surrogate model and acquisition function behind sample-efficient Bayesian optimization.
The Optuna framework, its TPE search algorithm, and the compute savings from trial pruning.
Defining success metrics and feasibility before any modeling begins, with a checklist.
Common data sources, rough volume guidance, and licensing/ethical considerations.
The core exploration checklist for catching data issues before modeling.
Detecting corrupt files, duplicates and mislabeled examples, with practical code.
The critical rule of fitting normalization statistics on training data only.
Assembling a complete augmentation pipeline with correct operation order.
Practical splitting code, stratification, and k-fold cross-validation for smaller datasets.
A practical decision framework for choosing architecture based on data size and constraints.
The training stage checklist, including the valuable tiny-batch overfitting sanity check.
Choosing the right metrics for the task and comparing against baselines honestly.
When to tune in the project lifecycle, and how to budget compute across tuning stages.
The core error analysis process for finding systematic patterns in model mistakes.
Building a complete model artifact bundle with config, preprocessing and metadata.
A preview of the deployment stages, and batch vs real-time inference.
Why deployed models degrade silently, and a practical prediction-drift monitoring signal.
What gets serialized, and why native PyTorch format is not always deployment-ready.
How Python pickle works, and the real security risk of unpickling untrusted files.
Tracing vs scripting for converting PyTorch models to a Python-independent format.
Exporting PyTorch models to a framework-agnostic format for portable inference.
Wrapping a trained model in a real callable web API using FastAPI.
Production-grade API considerations: validation, error handling, and health checks.
Containerizing a model-serving application for reproducible deployment.
When GPU inference is worth it, and dynamic batching for efficient GPU utilization.
Common cloud deployment patterns and the cold-start tradeoff in serverless.
Deploying models with SageMaker and other AWS services, with working code.
Running inference on accumulated data on a schedule, and why it is often simpler.
Serving individual requests immediately, and why percentile latency matters most.
Quantization, pruning, distillation and compilation for faster production inference.
Formalizing the project lifecycle into automated, reproducible stages.
Systematically recording every training run so results stay comparable.
A practical open-source tool for experiment tracking, models, and the model registry.
Centralized management of model versions and deployment stages.
Tracking exactly which dataset version was used for a given training run.
Systematically tracking every trained model version and its lineage.
The layers of production ML monitoring, from data quality to system health.
Detecting shifts in input data distribution with a statistical test.
When the relationship between inputs and correct outputs changes over time.
The observed performance decline that data and concept drift produce.
Comparing a new candidate model against production on live traffic.
Decomposing and profiling where prediction request time actually goes.
Maximizing requests handled per second, and its tradeoff with latency.
Diagnosing why a GPU sits idle and fixing common data loading bottlenecks.
Gradient checkpointing and mixed precision for fitting larger models in memory.
Data and model parallelism for training across multiple GPUs, with code.
How dense vector representations power modern search, retrieval and agent systems.
Specialized storage for efficient approximate nearest neighbor search at scale.
Combining retrieval with generation to ground LLM answers in real documents.
Models that jointly process and reason across text, images, audio and video.
Connecting pretrained vision encoders and language models for image understanding.
LLM-based systems that plan, use tools and take multi-step actions toward a goal.
The mechanism that lets a language model invoke external tools and APIs.
The structured JSON schema mechanism behind reliable LLM tool use.
Why sparse expert routing powers the largest modern language models.
Overcoming quadratic attention cost and the lost-in-the-middle problem.
Models trained to generate extended reasoning steps before a final answer.
Speculative decoding, continuous batching and KV-caching for large model serving.
A practical multi-pass strategy for reading deep learning papers efficiently.
Systematically surveying existing research before contributing new work.
Why fair, strong baselines are essential for credible research comparisons.
What state-of-the-art claims actually mean, and the benchmark saturation problem.
What makes a good benchmark, and the risk of benchmark gaming.
Systematically isolating which components of a method actually matter.
Why deep learning has a reproducibility challenge, and practices that help.
Core principles for planning trustworthy deep learning experiments.
Distinguishing genuine effects from random training noise, with code.
How model expressiveness connects to the bias-variance tradeoff in research.
How to compute parameter count, and why it is an imperfect complexity proxy.
Measuring actual computational cost as a complement to parameter count.
How to measure and report inference latency rigorously in research papers.
Estimating training and inference memory footprint, with practical formulas.
Index and prep strategy for the full Deep Learning interview question section.
Foundational Deep Learning interview questions with fully explained answers.
CNN interview questions covering convolution, pooling and key architectures.
RNN and LSTM interview questions covering gating, BPTT and vanishing gradients.
Transformer interview questions covering self-attention and architecture.
Optimizer, learning rate and regularization interview questions explained.
PyTorch interview questions covering autograd and the training loop, with code.
LLM interview questions covering pretraining, fine-tuning, sampling and KV-cache.
Deployment interview questions covering serving, security and production concerns.
Open-ended, situational Deep Learning interview scenarios, worked through.
Index and usage guide for the Deep Learning practice problems section.
Build a perceptron and MLP from scratch, with manual backpropagation.
Manual convolution, output shape math, and building a real CNN classifier.
Implement an RNN cell manually and build an LSTM sentiment classifier.
Implement scaled dot-product attention, masking and multi-head attention.
Custom Datasets, training loops, debugging exercises and custom losses.
Implement gradient descent, momentum and Adam from scratch.
Compute metrics, ROC curves and cross-validation from scratch.
Index and approach guide for the twelve end-to-end Deep Learning projects.
Build a transfer-learning image classifier end to end, with full code.
Fine-tune a pretrained detector on a custom dataset, with full code.
Build a U-Net for pixel-level semantic segmentation, with full code.
Build an LSTM text classifier from raw text to a trained model.
Train a character-level language model and generate new text with it.
Build a document-grounded RAG chatbot with citations, with full code.
Build and train a DCGAN to generate new images from noise.
Implement a simplified DDPM diffusion model from first principles.
Build a mini-GPT decoder-only Transformer entirely from scratch.
Fine-tune a pretrained LLM with LoRA on a custom instruction dataset.
Build a CNN encoder + LSTM decoder image captioning model.
Deploy a trained model as a real, containerized REST API service.
Join Coding Hubs School of AI's AI Engineering Diploma course โ live mentorship, hands-on projects, and 100% placement support in Delhi NCR.
Enroll Now โ Free Demo Available
Insights on AI, Data Science, Full Stack & Career