Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #571

Hidden Units Tuning

Practical guidance on choosing layer width (the number of hidden units per layer) — the companion consideration to Network Depth Tuning.

Common Practical Patterns

PatternDescription
Powers of 264, 128, 256, 512, ... — aligns well with GPU memory/compute organization, similar to batch size conventions
Funnel shapeProgressively narrowing width toward the output (e.g. 512 → 256 → 128 → num_classes) — a common, reasonable default pattern for classification heads
Uniform widthEvery hidden layer the same size — simpler to reason about, common in Transformer blocks where \(d_{\text{model}}\) stays constant throughout the stack

Why the Funnel Pattern Is Common

Progressively narrowing width toward the output mirrors the general idea that a network should gradually compress its representation from raw input features down toward the (typically much smaller) number of output classes — similar in spirit to an autoencoder's bottleneck (see Autoencoders), though not a strict requirement architecturally.

Code — A Simple Width Sweep

results = {}
for hidden_size in [64, 128, 256, 512]:
    model = nn.Sequential(
        nn.Linear(input_dim, hidden_size), nn.ReLU(),
        nn.Linear(hidden_size, num_classes)
    )
    train(model, train_loader, epochs=20)
    val_acc = evaluate(model, val_loader)
    results[hidden_size] = val_acc
    print(f"hidden_size={hidden_size}: val_accuracy={val_acc:.4f}, params={sum(p.numel() for p in model.parameters())}")

The Width/Data-Size Relationship

Wider layers add more parameters, which increases overfitting risk on small datasets — directly connecting to the exact tradeoff already discussed for network capacity generally in Bias-Variance Tradeoff. As with depth, width should be scaled thoughtfully relative to available training data, not maximized indiscriminately.

Common Mistakes

  • Choosing hidden layer sizes arbitrarily without any relationship to input/output dimensionality or available data size — very wide layers on a small dataset are a common, easily avoidable overfitting risk.
  • Tuning width and depth completely independently, ignoring their joint interaction — the same total parameter count can be achieved via different depth/width combinations, and these aren't always interchangeable in terms of performance, as flagged in Transformer Blocks.

Interview Relevance

Q: "Why might a 'funnel' pattern (progressively narrowing hidden layer sizes) be a reasonable default for a classification network's head?" It reflects the general idea that a network should progressively compress its representation from the raw input's dimensionality down toward the much smaller number of output classes needed for the final prediction — a natural, gradual reduction rather than an abrupt jump from a wide hidden layer directly to a small output layer.

Practice Question

Why does increasing hidden layer width carry a similar overfitting risk consideration to increasing network depth?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →