Coding Now โ€“ Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
๐Ÿ“ž Call Now: +91 9667708830
All Subjects
Free Learning Resource

๐Ÿง  Deep Learning Notes

Learn Deep Learning from first principles โ€” math, neural networks, CNNs, sequence models, Transformers, LLMs, PyTorch and production deployment.

๐Ÿง 

Foundations of Deep Learning

AI vs ML vs DL, why deep learning works now, the DL workflow and its real limitations.

Beginner ~3 hrs
11 topics โ†’
๐Ÿ“

Linear Algebra for DL

Scalars to eigenvectors โ€” the matrix and tensor operations every neural network runs on.

Beginnerโ€“Intermediate ~3.5 hrs
14 topics โ†’
๐Ÿ“ˆ

Calculus for DL

Derivatives, gradients and computational graphs โ€” the machinery behind learning.

Beginnerโ€“Intermediate ~2.5 hrs
10 topics โ†’
๐ŸŽฒ

Probability & Statistics for DL

Distributions, likelihood, entropy and KL divergence โ€” the language of loss functions.

Intermediate ~3 hrs
13 topics โ†’
๐Ÿ”—

Neural Network Fundamentals

From the biological neuron to the MLP โ€” layers, weights, forward propagation and the full training flow.

Beginnerโ€“Intermediate ~4 hrs
16 topics โ†’
โšก

Activation Functions

Every major activation function โ€” formula, graph, derivative and where it's actually used.

Beginnerโ€“Intermediate ~3 hrs
16 topics โ†’
๐ŸŽฏ

Loss Functions

Regression, classification and advanced losses โ€” with formulas and worked numbers.

Intermediate ~2.5 hrs
12 topics โ†’
๐Ÿš€

Optimization & LR Scheduling

SGD to AdamW, and every learning-rate schedule that actually gets used in practice.

Intermediate ~3.5 hrs
17 topics โ†’
๐Ÿ”„

Backpropagation

The chain rule applied to a real network, with a full numerical worked example.

Intermediate ~2.5 hrs
10 topics โ†’
๐Ÿ‹๏ธ

Training Deep Networks

Epochs, checkpoints, early stopping โ€” and the underfitting/overfitting tradeoff.

Intermediate ~2.5 hrs
11 topics โ†’
๐Ÿ›ก๏ธ

Regularization

L1/L2, dropout, weight decay and augmentation compared side by side.

Intermediate ~1.5 hrs
7 topics โ†’
๐Ÿ“

Normalization Techniques

BatchNorm, LayerNorm, GroupNorm and RMSNorm โ€” formulas and when each is used.

Intermediate ~1.5 hrs
6 topics โ†’
โœ…

Evaluation Metrics

Classification, regression and modern DL metrics โ€” from confusion matrix to BLEU and mAP.

Intermediate ~3 hrs
16 topics โ†’
๐Ÿ–ผ๏ธ

CNN Fundamentals

Convolution, kernels, stride, padding and pooling โ€” with a manual convolution example.

Intermediate ~3.5 hrs
17 topics โ†’
๐Ÿ›๏ธ

CNN Architectures

LeNet to ConvNeXt โ€” the innovation, and the problem, behind every landmark CNN.

Intermediateโ€“Advanced ~3 hrs
11 topics โ†’
๐Ÿ‘๏ธ

Computer Vision

Detection and segmentation families โ€” R-CNN through YOLO, FCN through Mask R-CNN.

Advanced ~4 hrs
18 topics โ†’
๐Ÿ”

Recurrent Neural Networks

Sequential data, hidden state, unrolling, and why plain RNNs struggle with long sequences.

Intermediate ~2 hrs
11 topics โ†’
๐Ÿงฌ

LSTM & GRU

Gates, cell state and the full equations for LSTM and GRU, compared directly.

Intermediateโ€“Advanced ~3 hrs
17 topics โ†’
๐Ÿ”

Seq2Seq & Attention

Encoder-decoder, the context-vector bottleneck, and attention built from first principles (Q/K/V).

Advanced ~2.5 hrs
13 topics โ†’
๐Ÿงฉ

Transformers

The full architecture โ€” encoder, decoder, positional encoding and residual connections โ€” with data flow.

Advanced ~3 hrs
10 topics โ†’
๐Ÿ“

NLP with Deep Learning

Tokenization to embeddings โ€” Word2Vec, GloVe, BERT, RoBERTa, T5 and GPT-style models.

Intermediateโ€“Advanced ~3 hrs
14 topics โ†’
๐Ÿค–

LLM Fundamentals

Pretraining to RLHF/DPO, context windows, sampling and KV-cache โ€” how LLMs actually run.

Advanced ~4 hrs
19 topics โ†’
๐ŸŽจ

Generative Deep Learning

Autoencoders, VAEs and GANs โ€” generative vs discriminative, explained from the ground up.

Advanced ~2.5 hrs
9 topics โ†’
๐ŸŒซ๏ธ

Diffusion Models

Forward noising, reverse denoising, latent diffusion and Stable Diffusion.

Advanced ~2.5 hrs
10 topics โ†’
๐Ÿ”€

Transfer Learning

Pretrained models, freezing layers, and fine-tuning vs feature extraction.

Intermediate ~1.5 hrs
7 topics โ†’
๐ŸŽ›๏ธ

Modern Fine-Tuning (PEFT)

LoRA, QLoRA, adapters and quantization โ€” with the LoRA math worked out.

Advanced ~2 hrs
10 topics โ†’
๐Ÿ”ฎ

Self-Supervised Learning

Pretext tasks, contrastive learning (SimCLR, MoCo) and masked modeling.

Advanced ~1.5 hrs
8 topics โ†’
๐Ÿงช

Advanced Deep Learning

Few-shot, meta-learning, distillation, federated learning and Mixture of Experts.

Advanced ~2 hrs
10 topics โ†’
๐Ÿ”ฅ

PyTorch

Tensors, autograd, nn.Module, DataLoader and full training loops โ€” hands-on.

Intermediate ~4.5 hrs
21 topics โ†’
๐Ÿ“Š

TensorFlow & Keras

Sequential and Functional APIs, callbacks, and how TensorFlow compares to PyTorch.

Intermediate ~2 hrs
11 topics โ†’
๐ŸŽš๏ธ

Hyperparameter Tuning

What to tune, and how โ€” grid search, random search, Bayesian optimization and Optuna.

Intermediateโ€“Advanced ~2 hrs
14 topics โ†’
๐Ÿ—๏ธ

DL Project Development

The full lifecycle โ€” problem definition through monitoring โ€” applied end to end.

Intermediateโ€“Advanced ~2.5 hrs
15 topics โ†’
๐Ÿ“ฆ

Deployment

TorchScript, ONNX, FastAPI, Docker and cloud deployment for real inference traffic.

Advanced ~2.5 hrs
13 topics โ†’
โš™๏ธ

Production DL & MLOps

Experiment tracking, model registries, drift detection and distributed training.

Advanced ~3 hrs
16 topics โ†’
โœจ

Modern AI

Embeddings, vector databases, RAG, agents and long-context/reasoning models โ€” from a DL lens.

Advanced ~2.5 hrs
12 topics โ†’
๐Ÿ“š

Research Concepts

Reading papers, baselines, SOTA, ablations, reproducibility, FLOPs and benchmarking.

Advanced ~2 hrs
14 topics โ†’
๐ŸŽฏ

Interview Questions

Topic-wise Deep Learning interview questions with fully explained answers.

All levels ~3.5 hrs
10 topics โ†’
๐Ÿ“

Practice Questions

Implementation-oriented exercises across networks, CNNs, RNNs, transformers and PyTorch.

All levels ~3 hrs
8 topics โ†’
๐Ÿ’ผ

Deep Learning Projects

End-to-end builds โ€” image classifiers through fine-tuned LLMs โ€” with deployment.

Intermediateโ€“Advanced ~6 hrs
13 topics โ†’

No notes found. Try a different search term, or browse all Deep Learning notes.

๐Ÿง  Foundations of Deep Learning

11 of 11 topics published
101

What Is Deep Learning?

What Artificial Intelligence is, its sub-fields, and how deep learning fits inside it โ€” the map for this whole notes hub.

102

AI vs ML vs Deep Learning

The exact containment relationship between AI, Machine Learning and Deep Learning, with a side-by-side comparison table.

103

Traditional ML vs Deep Learning

Manual feature engineering vs automatic feature learning, compared side by side with a pipeline diagram.

104

Evolution of Deep Learning

From the 1958 Perceptron through two AI winters to the 2012 breakthrough and the modern LLM era.

105

Why Deep Learning Became Successful

The three pillars โ€” data, compute and algorithmic fixes โ€” that made deep networks finally trainable at scale.

106

Applications of Deep Learning

Real-world deep learning applications across computer vision, NLP, healthcare, finance and generative AI.

107

Types of Learning in Deep Learning

Supervised, unsupervised, semi-supervised, self-supervised and reinforcement learning in a deep learning context.

108

Deep Learning Workflow

The end-to-end deep learning project workflow, from problem definition to deployment and monitoring.

109

Components of a Deep Learning System

Data, model, loss function, optimizer and hardware โ€” the components every trained deep learning system depends on.

110

Challenges in Deep Learning

Data hunger, compute cost, overfitting, vanishing gradients and interpretability โ€” the real engineering challenges.

111

Advantages & Limitations of Deep Learning

A practical decision guide for when deep learning's advantages outweigh its costs, and when they don't.

๐Ÿ“ Linear Algebra for DL

14 of 14 topics published
201

Scalars, Vectors & Matrices

The three basic mathematical objects behind every neural network โ€” scalars, vectors and matrices โ€” with notation and shapes explained.

202

Tensors

Tensors as the n-dimensional generalization of scalars, vectors and matrices, with rank, shape and PyTorch code.

203

Matrix Addition

Element-wise matrix addition, its formula, and its role in bias terms and ResNet residual connections.

204

Matrix Multiplication

The shape rule, a full worked example, and why matrix multiplication is the core operation of every linear layer.

205

Dot Product

The dot product formula, geometric intuition, and its role in cosine similarity and attention scores.

206

Matrix Transpose

What matrix transpose does, its key properties, and its role in backpropagation and attention.

207

Matrix Inverse

The matrix inverse, the 2x2 formula, and why deep learning uses iterative optimization instead of computing inverses.

208

Determinant

The determinant formula, its geometric meaning as area/volume scaling, and its link to invertibility.

209

Vector Norms

L1, L2 and L-infinity norms compared, with their role in regularization and gradient clipping.

210

Eigenvalues

What eigenvalues are, the characteristic equation, a worked numerical example, and their role in PCA.

211

Eigenvectors

What eigenvectors are, how to find them, and how they power Principal Component Analysis.

212

Vector Spaces

The formal definition of a vector space, basis, span and dimension, connected to feature and embedding spaces.

213

Linear Transformations

What linear transformations are, common examples, and why they motivate the need for activation functions.

214

Broadcasting

The broadcasting rule in NumPy and PyTorch, worked examples, and common shape bugs to avoid.

๐Ÿ“ˆ Calculus for DL

10 of 10 topics published

๐ŸŽฒ Probability & Statistics for DL

13 of 13 topics published
225

Probability Basics

Core probability definitions and axioms, and why a classifier's softmax output must satisfy them.

226

Random Variables

Discrete vs continuous random variables, and how dataset labels and weight initialization use them.

227

Probability Distributions

PMFs, PDFs, the Bernoulli, categorical and normal distributions, and their roles across deep learning.

228

Expected Value

The formula for expected value, and how training loss is literally an empirical expected value.

229

Variance & Standard Deviation

Variance and standard deviation formulas, and why they motivate feature standardization and batch normalization.

230

Covariance

The covariance formula, the covariance matrix, and its role in PCA and dimensionality reduction.

231

Conditional Probability

The conditional probability formula, a spam-filter worked example, and independence.

232

Bayes' Theorem

Bayes' theorem formula, prior/likelihood/posterior, and a full worked medical-test example.

233

Likelihood

The difference between probability and likelihood, with a coin-flip worked example.

234

Maximum Likelihood Estimation

MLE explained, and a full derivation of cross-entropy and MSE loss from maximum likelihood principles.

235

Entropy

The entropy formula, worked examples, and how it measures uncertainty in a probability distribution.

236

Cross-Entropy

The cross-entropy formula, a full worked example, and PyTorch's CrossEntropyLoss explained.

237

KL Divergence

The KL divergence formula, its relationship to cross-entropy and entropy, and its use in VAEs and RLHF.

๐Ÿ”— Neural Network Fundamentals

16 of 16 topics published
238

Biological Neuron

The biological neuron structure that inspired artificial neural networks, mapped part by part.

239

Artificial Neuron

The full artificial neuron computation โ€” weighted sum plus activation โ€” with a worked numerical example.

240

McCulloch-Pitts Neuron

The first mathematical neuron model from 1943, and how it modeled logical AND and OR without learning.

241

Perceptron

The 1958 Perceptron model, its linear decision boundary, and what makes it different from earlier neuron models.

242

Perceptron Learning Algorithm

The Perceptron's weight update rule, a full worked example, and the Perceptron Convergence Theorem.

243

Limitations of Perceptron

Why a single Perceptron cannot learn XOR, proven algebraically, and how a hidden layer fixes it.

244

Multi-Layer Perceptron

The MLP architecture that solves XOR, with a diagram and full PyTorch implementation.

245

Neural Network Architecture

Input, hidden and output layers, depth vs width, and common architecture patterns.

246

Input, Hidden & Output Layers

What each layer type does, and how to correctly size and activate the output layer for any task.

247

Weights and Bias

What weights and bias represent, weight matrix shapes, and why zero-initialization fails.

248

Parameters vs Hyperparameters

The precise distinction between learned parameters and chosen hyperparameters, with clear examples.

249

Forward Propagation

The forward propagation formula, a full numerical walkthrough, and matching PyTorch code.

250

Loss Function Basics

What a loss function measures, MSE vs cross-entropy introduced, and loss vs evaluation metric.

251

Backpropagation (Intro)

A conceptual introduction to backpropagation as efficient gradient computation via the chain rule.

252

Gradient Descent (Intro)

The gradient descent update rule, why learning rate matters, and a complete PyTorch training step.

253

Weight Updates

A complete worked example of one full training step, from forward pass to weight update, verified in PyTorch.

โšก Activation Functions

16 of 16 topics published
254

Step Function

The original 1958 Perceptron activation function, and why its zero gradient makes it unusable for backpropagation.

255

Linear Activation

The linear (identity) activation, why it collapses network depth, and where it's actually used in regression outputs.

256

Sigmoid Function

The sigmoid formula, derivative, graph, and why it causes vanishing gradients in deep networks.

257

Tanh Function

The tanh formula, derivative, and how it improves on sigmoid while sharing its saturation problem.

258

ReLU

The ReLU formula, why it fixes vanishing gradients, and the dying ReLU problem explained.

259

Leaky ReLU

How Leaky ReLU's small negative slope fixes the dying ReLU problem, with formula and code.

260

PReLU

How PReLU makes the negative slope a learnable parameter instead of a fixed hyperparameter.

261

ELU

The ELU formula, its smooth exponential negative curve, and its tradeoffs versus ReLU.

262

SELU

The self-normalizing SELU activation, its strict requirements, and why it's used less often in practice.

263

GELU

The GELU formula, its probabilistic derivation, and why it's the default in Transformer architectures.

264

Swish

The Swish activation function, discovered via search, and its close relationship to GELU.

265

Softmax Function

The softmax formula, a full worked numerical example, and the numerical stability trick every framework uses.

266

Sigmoid vs Tanh

A direct comparison of sigmoid and tanh, and why both still appear inside LSTM and GRU gates.

267

ReLU vs Leaky ReLU

A direct comparison of ReLU and Leaky ReLU, and when the dying ReLU problem actually justifies switching.

268

ReLU vs GELU

Why CNNs favor ReLU while Transformers favor GELU, compared side by side.

269

Sigmoid vs Softmax

When to use sigmoid vs softmax for an output layer, based on whether classes are mutually exclusive.

๐ŸŽฏ Loss Functions

12 of 12 topics published

๐Ÿš€ Optimization & LR Scheduling

17 of 17 topics published
282

Gradient Descent

The optimization problem formalized, convex vs non-convex loss surfaces, and iterative descent visualized.

283

Batch Gradient Descent

How batch gradient descent uses the full dataset per update, and why it's rarely used in practice.

284

Stochastic Gradient Descent

How SGD uses one random example per update, its noisy path, and why that noise can help.

285

Mini-Batch Gradient Descent

The practical compromise between batch and stochastic gradient descent, and how it maps to PyTorch's DataLoader.

286

Learning Rate

A deep dive on the learning rate hyperparameter, typical ranges, and its interaction with batch size.

287

Momentum

The momentum formula, the physical rolling-ball analogy, and how it smooths oscillation in narrow ravines.

288

Nesterov Momentum

How Nesterov momentum's look-ahead gradient improves on standard momentum.

289

AdaGrad

The AdaGrad formula, per-parameter adaptive learning rates, and its diminishing learning rate flaw.

290

RMSProp

How RMSProp fixes AdaGrad's diminishing learning rate problem with a decaying average.

291

Adam Optimizer

The Adam optimizer's formula, bias correction explained, and why it became the default optimizer.

292

AdamW

How AdamW fixes Adam's broken interaction with L2 regularization, with a full optimizer comparison table.

293

Learning Rate Scheduling

Why a fixed learning rate is a compromise, and an overview of the major scheduling strategies.

294

Step Decay

The step decay formula, a worked example, and its staircase-shaped learning rate curve.

295

Exponential Decay

The exponential decay formula and how it compares to step decay's sudden drops.

296

Cosine Annealing

The cosine annealing formula, why its smooth shape is popular, and warm restarts explained.

297

Warmup Learning Rate

Why gradually increasing the learning rate at the start of training stabilizes early updates.

298

One Cycle Learning Rate

The One Cycle policy's rise-fall-anneal shape and how it enables faster "super-convergence" training.

๐Ÿ”„ Backpropagation

10 of 10 topics published

๐Ÿ‹๏ธ Training Deep Networks

11 of 11 topics published

๐Ÿ›ก๏ธ Regularization

7 of 7 topics published

๐Ÿ“ Normalization Techniques

6 of 6 topics published

โœ… Evaluation Metrics

16 of 16 topics published
333

Confusion Matrix

The 2x2 confusion matrix foundation โ€” TP, TN, FP, FN โ€” every classification metric is built from.

334

Accuracy

The accuracy formula and why it can be deeply misleading on imbalanced classification datasets.

335

Precision & Recall

The precision and recall formulas, their tradeoff, and how to choose which matters more by context.

336

F1 Score

The F1 score formula, why it uses a harmonic mean, and the F-beta generalization.

337

Specificity

The specificity (true negative rate) formula and how it complements recall.

338

ROC Curve

The ROC curve's TPR vs FPR plot across every threshold, with a worked code example.

339

ROC-AUC

The ROC-AUC metric, its probabilistic interpretation, and a worked numerical example.

340

Precision-Recall Curve

The Precision-Recall curve and why it's more informative than ROC for imbalanced datasets.

341

PR-AUC

The PR-AUC metric, its class-imbalance-dependent baseline, and when to prefer it over ROC-AUC.

342

Regression Metrics

The Rยฒ metric formula, a worked example, and how MAE/MSE/RMSE double as evaluation metrics.

343

Perplexity

The perplexity formula for language model evaluation, and its intuitive branching-factor interpretation.

344

BLEU Score

The BLEU score formula, n-gram precision, the brevity penalty, and its known limitations.

345

ROUGE Score

The ROUGE score formula, why it emphasizes recall, and how it compares to BLEU.

346

IoU (Intersection over Union)

The IoU formula, a worked bounding box example, and its role in object detection matching.

347

Dice Score

The Dice score formula, its exact relationship to IoU, and why it's popular in medical imaging.

348

Mean Average Precision

How mAP combines IoU matching and per-class precision-recall into the standard object detection metric.

๐Ÿ–ผ๏ธ CNN Fundamentals

17 of 17 topics published
349

What Is CNN?

A high-level introduction to CNNs, the convolution/pooling/FC pipeline, and a minimal PyTorch example.

350

Why CNN?

Why plain MLPs fail on images, and how sparse connectivity and parameter sharing fix it, with a parameter comparison.

351

Image Representation

How images are represented as tensors, pixel value ranges, and channel-first vs channel-last conventions.

352

Convolution Operation

The convolution formula, a worked numerical example, and its implementation in PyTorch.

353

Kernel and Filter

The precise difference between a kernel and a filter, with classic edge-detection kernel examples.

354

Feature Map

What a feature map represents, and how multiple filters produce multiple stacked feature maps.

355

Stride

The stride parameter, its output size formula, and why larger strides act as built-in downsampling.

356

Padding

The padding parameter, valid vs same padding, and why it prevents edge-pixel underprocessing.

357

Channels

How multi-channel convolution combines input channels, and input vs output channel counts.

358

Receptive Field

How receptive field grows across stacked layers, and why small kernels are stacked instead of large ones.

359

Pooling

The general pooling concept, why it helps, and its zero-parameter downsampling benefit.

360

Max Pooling

The max pooling formula, a worked numerical example, and why it suits feature detection.

361

Average Pooling

The average pooling formula, a worked example, and when it's preferred over max pooling.

362

Global Average Pooling

How GAP collapses entire feature maps, with a dramatic parameter-count comparison to flattening.

363

Flattening

How flattening bridges spatial feature maps to fully connected layers, with a worked example.

364

Fully Connected Layer

The role of the fully connected layer at the end of a CNN, and why it often dominates parameter count.

365

Manual Convolution Example

A complete manual convolution example with padding and stride, verified against PyTorch's conv2d.

๐Ÿ›๏ธ CNN Architectures

11 of 11 topics published

๐Ÿ‘๏ธ Computer Vision

18 of 19 topics published
377

Image Classification

The core image classification task, the standard CNN pipeline, and benchmark datasets.

378

Transfer Learning

Feature extraction vs fine-tuning for reusing pretrained CNNs, with PyTorch code for both.

379

Object Detection

The object detection task, one-stage vs two-stage detector families, and evaluation with mAP.

380

Image Segmentation

The general pixel-level segmentation task, and the three flavors: semantic, instance and panoptic.

381

Semantic Segmentation

Per-pixel class labeling without instance distinction, with a worked numerical example.

382

Instance Segmentation

How instance segmentation distinguishes individual objects, built via detection plus masking.

383

Panoptic Segmentation

How panoptic segmentation unifies "things" and "stuff" into one complete scene output.

384

Face Recognition

Embedding-based face verification and identification using triplet and contrastive loss.

385

OCR

The two-stage OCR pipeline of text detection and sequence-based text recognition.

386

Image Captioning

The CNN encoder plus RNN/Transformer decoder pipeline for generating image descriptions.

387

Pose Estimation

Heatmap-based keypoint detection for reconstructing body pose from an image.

388

R-CNN

The original region-proposal-based object detector, and why it was extremely slow.

389

Fast R-CNN

How Fast R-CNN shares CNN computation across regions using ROI Pooling.

390

Faster R-CNN

How the Region Proposal Network made object detection fully end-to-end trainable.

391

SSD (Object Detection)

How SSD predicts detections in a single pass using multi-scale feature maps.

392

YOLO

The grid-based single-pass detection approach behind YOLO's real-time speed.

393

FCN (Fully Convolutional Network)

How FCN replaced fully connected layers with convolutions for practical semantic segmentation.

395

Mask R-CNN

How Mask R-CNN adds instance segmentation to Faster R-CNN with ROI Align and a mask branch.

๐Ÿ” Recurrent Neural Networks

11 of 11 topics published

๐Ÿงฌ LSTM & GRU

17 of 17 topics published
407

Why LSTM

The core motivation for LSTM โ€” a cell state pathway that protects gradients across long sequences.

408

LSTM Architecture

The five components of an LSTM cell and how they connect, with a diagram and PyTorch weight shapes.

409

LSTM Cell State

The cell state update formula and why its largely-additive structure protects gradients across time.

410

LSTM Hidden State

How the LSTM hidden state is derived from the cell state, and why the two states are kept separate.

411

Forget Gate

The LSTM forget gate formula, a worked example, and how it decides what memory to discard.

412

Input Gate

The LSTM input gate formula, a worked example, and how it decides how much new information to add.

413

Candidate State

The LSTM candidate state formula, why it uses tanh instead of sigmoid, and a worked example.

414

Output Gate

The LSTM output gate formula and how it controls what the cell state exposes as the hidden state.

415

LSTM Equations

All six LSTM equations assembled into one complete reference, with a summary table and diagram.

416

LSTM Forward Pass

A complete numerical LSTM forward pass through one time step, verified against PyTorch.

417

LSTM Advantages & Limitations

LSTM's genuine strengths and remaining weaknesses, including the parallelization problem it never solved.

418

LSTM Applications

Real-world LSTM applications โ€” translation, speech recognition, forecasting โ€” with PyTorch code examples.

419

GRU Architecture

GRU's simplified structure compared to LSTM โ€” one state, two gates instead of two states, three gates.

420

GRU Update Gate

The GRU update gate formula, how it merges LSTM's forget and input gates, with a worked example.

421

GRU Reset Gate

The GRU reset gate formula and how it shapes the candidate hidden state computation.

422

GRU Equations

All four GRU equations assembled together, with a complete numerical worked example.

423

GRU vs LSTM

A complete side-by-side comparison of GRU and LSTM, with practical guidance on which to choose.

๐Ÿ” Seq2Seq & Attention

13 of 13 topics published
424

Encoder-Decoder Architecture

How splitting a model into an encoder and decoder decouples input and output sequence lengths.

425

Context Vector

The fixed-size context vector, why it's a bottleneck, and its role in basic Seq2Seq models.

426

Seq2Seq Model

The complete assembled Seq2Seq pipeline and autoregressive decoding, with a full code example.

427

Teacher Forcing

How teacher forcing speeds up Seq2Seq training, and the exposure bias it introduces.

428

Seq2Seq Limitations

The complete list of basic Seq2Seq limitations that directly motivated the attention mechanism.

429

Why Attention

The core motivation for attention, and how it directly fixes the Seq2Seq context-vector bottleneck.

430

Query, Key, Value

The Q/K/V framework explained via a search-engine analogy, with formulas and a numerical example.

431

Attention Score

How the dot product between query and key produces a raw relevance score, with a worked example.

432

Dot-Product Attention

The complete dot-product attention formula, from scores to softmax to weighted value sum.

433

Scaled Dot-Product Attention

Why attention scores are divided by the square root of dk, and how this keeps softmax gradients healthy.

434

Self-Attention

How self-attention relates every token in a sequence directly, without an RNN's distance penalty.

435

Cross-Attention

How cross-attention lets a decoder query an encoder's outputs, formally solving the Seq2Seq bottleneck.

436

Multi-Head Attention

How running multiple attention heads in parallel lets a model capture different relationship types.

๐Ÿงฉ Transformers

10 of 10 topics published

๐Ÿ“ NLP with Deep Learning

14 of 14 topics published
447

Text Preprocessing

Classical text preprocessing steps and why modern Transformers rely on much less of them.

448

Tokenization

Word, character and subword tokenization compared, and why BPE became the modern standard.

449

Vocabulary

The vocabulary size tradeoff and how it directly sizes a model's embedding and output layers.

450

One-Hot Encoding (NLP)

Why one-hot word encoding fails โ€” huge dimensionality and zero notion of similarity.

451

Word Embeddings

How dense word embeddings capture semantic similarity, with the famous king-queen analogy.

452

Word2Vec

The self-supervised idea behind Word2Vec, and a preview of its CBOW and Skip-Gram architectures.

453

CBOW

The Continuous Bag of Words formula and how it predicts a target word from averaged context.

454

Skip-Gram

The Skip-Gram formula, and a direct comparison with CBOW including rare-word performance.

455

GloVe

How GloVe learns embeddings from global co-occurrence statistics, compared to Word2Vec.

456

Contextual Embeddings

Why static embeddings can't distinguish word meanings, and how contextual embeddings fix this.

457

BERT

BERT's bidirectional architecture, Masked Language Modeling, and Next Sentence Prediction explained.

458

RoBERTa

How RoBERTa improved on BERT using the identical architecture, purely through a better training recipe.

459

T5 Model

How T5 reframes every NLP task as text-to-text using a full encoder-decoder Transformer.

460

GPT Architecture

GPT's decoder-only architecture, next-token prediction, and a comparison against BERT.

๐Ÿค– LLM Fundamentals

19 of 19 topics published
461

What Is an LLM?

What makes a model "large," emergent capabilities, and why LLMs aren't automatically chatbots.

462

LLM Architecture

The modern refinements (RMSNorm, RoPE, grouped-query attention) on top of the base Transformer decoder.

463

Tokens and Tokenization (LLMs)

Why token counts matter for cost, context budget and latency, with tiktoken code.

464

Embeddings (LLMs)

The LLM input embedding layer at scale, and the weight-tying trick between input and output layers.

465

Positional Embeddings

Learned positional embeddings vs Rotary Positional Embeddings (RoPE), and why RoPE became dominant.

466

Transformer Blocks

The transformer block as the fundamental unit of LLM scale, and depth vs width tradeoffs.

467

LLM Pretraining

The massive self-supervised pretraining stage, and a brief look at scaling laws.

468

Next-Token Prediction

The full formula for next-token prediction and the parallel-supervision trick causal masking enables.

469

Supervised Fine-Tuning

How SFT reshapes a pretrained model using a much smaller, curated dataset of example responses.

470

Instruction Tuning

How training on diverse instruction-phrased tasks teaches models to generalize to unseen instructions.

471

LLM Alignment

Why capability and alignment are distinct concerns, and what alignment actually aims to achieve.

472

RLHF

The full 3-stage RLHF process โ€” reward model training and RL fine-tuning with a KL penalty.

473

DPO

How Direct Preference Optimization achieves RLHF-like results without a reward model or RL.

474

Context Window

Why attention's quadratic cost limits context window size, with a practical token-budget example.

475

Temperature (Sampling)

The temperature sampling formula and how it sharpens or flattens a model's output distribution.

476

Top-K Sampling

The top-K sampling algorithm, a worked example, and its fixed-count limitation.

477

Top-P Sampling

Nucleus (top-P) sampling, how it adapts to distribution shape, compared directly to top-K.

478

KV Cache

Why the KV cache is valid, and how it turns quadratic generation cost into linear cost.

479

LLM Perplexity

The perplexity formula, a worked example, and what perplexity does and doesn't measure.

๐ŸŽจ Generative Deep Learning

9 of 9 topics published

๐ŸŒซ๏ธ Diffusion Models

10 of 10 topics published

๐Ÿ”€ Transfer Learning

7 of 7 topics published

๐ŸŽ›๏ธ Modern Fine-Tuning (PEFT)

10 of 10 topics published

๐Ÿ”ฎ Self-Supervised Learning

8 of 8 topics published

๐Ÿงช Advanced Deep Learning

10 of 10 topics published

๐Ÿ”ฅ PyTorch

21 of 21 topics published
534

PyTorch Installation

Installing PyTorch with GPU support and verifying CUDA availability correctly.

535

PyTorch Tensors

Practical tensor creation, inspection, and the NumPy conversion memory-sharing gotcha.

536

Tensor Operations

Indexing, reshaping, view vs reshape, and the cat vs stack distinction, with code.

537

Broadcasting (PyTorch)

The classic silent shape-mismatch bug in real PyTorch code, with a worked debugging example.

538

Autograd

The practical autograd API โ€” requires_grad, backward(), no_grad(), and gradient accumulation.

539

Computational Graphs (PyTorch)

How PyTorch builds dynamic graphs on every forward pass, and retain_graph explained.

540

nn.Module

The base class for all PyTorch models, parameter registration, and train/eval modes.

541

PyTorch Layers

A practical reference catalog of common PyTorch layer types with exact syntax.

542

PyTorch Activations

Module vs functional activation forms, with a full quick-reference table.

543

PyTorch Loss Functions

A practical reference for PyTorch loss classes and their exact input format requirements.

544

PyTorch Optimizers

A practical reference for PyTorch optimizers, including per-parameter-group learning rates.

545

PyTorch Dataset

The minimal Dataset interface โ€” __len__ and __getitem__ โ€” with a full working example.

546

PyTorch DataLoader

DataLoader parameters, batching, shuffling, and custom collate functions for variable-length data.

547

PyTorch Training Loop

The complete practical training loop with device management and memory-safe loss logging.

548

PyTorch Validation Loop

The complete practical validation loop with correctly weighted loss averaging.

549

PyTorch GPU/CUDA

Moving models and data to GPU correctly, device-mismatch errors, and mixed precision training.

550

Saving PyTorch Models

Practical syntax for saving weights, full checkpoints, and best-model tracking.

551

Loading PyTorch Models

Practical syntax for loading weights, checkpoints, and partial weights across devices.

552

PyTorch Transfer Learning

Complete, runnable PyTorch code for feature extraction and partial fine-tuning.

553

PyTorch Custom Datasets

A complete realistic custom Dataset loading images from disk with CSV labels.

554

PyTorch Custom Training Loops

When to write a custom training loop, with gradient accumulation and custom logging patterns.

๐Ÿ“Š TensorFlow & Keras

11 of 11 topics published

๐ŸŽš๏ธ Hyperparameter Tuning

14 of 14 topics published
566

Hyperparameters Overview

A complete catalog of every hyperparameter covered in this hub, and the general tuning philosophy.

567

Learning Rate Tuning

The learning rate range test technique and how to diagnose too-high vs too-low learning rates.

568

Batch Size Tuning

The linear scaling rule, memory constraints, and the generalization tradeoff of batch size.

569

Epochs Tuning

Why modern practice relies on early stopping rather than fixing epoch count manually.

570

Network Depth Tuning

A practical strategy for choosing network depth based on underfitting/overfitting symptoms.

571

Hidden Units Tuning

Common width patterns like the funnel shape, and the width/data-size relationship.

572

Dropout Tuning

Typical dropout rates, where to apply dropout, and diagnosing rate adjustments from symptoms.

573

Weight Decay Tuning

Typical weight decay ranges and why AdamW is essential for meaningful weight decay tuning.

574

Optimizer Selection

A practical decision guide for choosing between AdamW, SGD and RMSProp.

575

Activation Function Selection

A practical decision guide for choosing activation functions by layer type and architecture.

576

Grid Search

The exhaustive grid search algorithm, its combinatorial explosion problem, and code.

577

Random Search

Why random search often outperforms grid search for the same compute budget.

578

Bayesian Optimization

The surrogate model and acquisition function behind sample-efficient Bayesian optimization.

579

Optuna

The Optuna framework, its TPE search algorithm, and the compute savings from trial pruning.

๐Ÿ—๏ธ DL Project Development

15 of 15 topics published
580

DL Problem Definition

Defining success metrics and feasibility before any modeling begins, with a checklist.

581

Dataset Collection

Common data sources, rough volume guidance, and licensing/ethical considerations.

582

Data Exploration

The core exploration checklist for catching data issues before modeling.

583

Data Cleaning

Detecting corrupt files, duplicates and mislabeled examples, with practical code.

584

Data Preprocessing

The critical rule of fitting normalization statistics on training data only.

585

Data Augmentation Pipeline

Assembling a complete augmentation pipeline with correct operation order.

586

Train/Val/Test Split (Lifecycle)

Practical splitting code, stratification, and k-fold cross-validation for smaller datasets.

587

Model Selection

A practical decision framework for choosing architecture based on data size and constraints.

588

Model Training

The training stage checklist, including the valuable tiny-batch overfitting sanity check.

589

Model Evaluation

Choosing the right metrics for the task and comparing against baselines honestly.

590

Hyperparameter Tuning (Lifecycle)

When to tune in the project lifecycle, and how to budget compute across tuning stages.

591

Error Analysis

The core error analysis process for finding systematic patterns in model mistakes.

592

DL Model Saving (Lifecycle)

Building a complete model artifact bundle with config, preprocessing and metadata.

593

DL Deployment Lifecycle

A preview of the deployment stages, and batch vs real-time inference.

594

DL Monitoring

Why deployed models degrade silently, and a practical prediction-drift monitoring signal.

๐Ÿ“ฆ Deployment

13 of 13 topics published

โš™๏ธ Production DL & MLOps

16 of 16 topics published

โœจ Modern AI

12 of 12 topics published

๐Ÿ“š Research Concepts

14 of 14 topics published

๐ŸŽฏ Interview Questions

10 of 10 topics published

๐Ÿ“ Practice Questions

8 of 8 topics published

๐Ÿ’ผ Deep Learning Projects

13 of 13 topics published

Ready to go from notes to a real career?

Join Coding Hubs School of AI's AI Engineering Diploma course โ€” live mentorship, hands-on projects, and 100% placement support in Delhi NCR.

Enroll Now โ€” Free Demo Available
๐Ÿ’ฌ Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles โ†’