Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #575

Activation Function Selection

Practical guidance for choosing which activation function to use where — synthesizing the comparison notes from the Activation Functions category into direct, actionable recommendations.

The Practical Defaults by Layer Type and Architecture

WhereDefault ChoiceWhy
Hidden layers, CNNs/MLPsReLUCheap, non-saturating for positive inputs, well-understood, strong empirical track record — see ReLU
Hidden layers, TransformersGELUEmpirically favored in this specific regime, standard in BERT/GPT-family feed-forward blocks — see ReLU vs GELU
Binary classification outputSigmoidDirectly produces \(P(\text{class}=1)\) — see Sigmoid Function
Multi-class classification outputSoftmaxProduces a valid probability distribution over mutually exclusive classes — see Softmax Function
Regression outputNone (linear/identity)Allows unbounded output values — see Linear Activation
LSTM/GRU gatesSigmoid (gates), Tanh (candidate states)Bounded ranges carry specific semantic meaning here — see Sigmoid vs Tanh

When to Deviate From ReLU in Hidden Layers

If training diagnostics reveal a significant fraction of "dead" neurons (see the dying ReLU problem from ReLU), trying Leaky ReLU or ELU (see ReLU vs Leaky ReLU) is a reasonable, targeted adjustment rather than a default choice made preemptively without evidence of the specific problem.

Code — Diagnosing Dead ReLU Neurons

import torch

def check_dead_relus(model, val_loader):
    activation_counts = {}
    hooks = []

    def make_hook(name):
        def hook(module, input, output):
            activation_counts[name] = activation_counts.get(name, 0) + (output > 0).float().mean().item()
        return hook

    for name, module in model.named_modules():
        if isinstance(module, torch.nn.ReLU):
            hooks.append(module.register_forward_hook(make_hook(name)))

    for x, _ in val_loader:
        model(x)
        break   # a single batch is often enough for a quick diagnostic

    for h in hooks:
        h.remove()
    return activation_counts   # low values suggest a large fraction of "dead" (always-zero) activations

Common Mistakes

  • Reflexively using a variant like Leaky ReLU or GELU everywhere "just in case," without evidence the plain default is actually causing a problem — this adds unnecessary complexity and (for GELU specifically) compute cost without a clear justification.
  • Using sigmoid or tanh for hidden layers in a deep feedforward network by default — as covered extensively in the Activation Functions category, this reintroduces vanishing gradients that ReLU-family activations largely resolved.

Interview Relevance

Q: "Why would you choose GELU for a Transformer-based model's hidden layers but ReLU for a CNN's hidden layers?" This reflects an empirical, architecture-specific finding rather than a universal ranking — GELU's smooth, non-monotonic shape has been found to help optimization stability specifically in the deep, large-scale, self-attention-heavy regime of Transformers (established starting with BERT), while CNNs, more sensitive to per-operation compute cost applied at every spatial location, generally haven't shown a strong enough benefit from GELU to justify replacing ReLU's simplicity and speed.

Practice Question

You're building the output layer for a model predicting a house's price (a continuous, unbounded positive value). What activation function (if any) would you use, and why?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →