Computing Library › Neural Architectures
Neural Architectures

Weight Initialization

How a network's weights are set before training determines whether signals and gradients stay at a usable scale through many layers.

Why initialization matters

Deep networks are trained by gradient descent from a starting point, and that starting point matters more than it might seem. If initial weights are too large, activations and gradients grow across layers and explode; too small, and they shrink toward zero and vanish. Poor initialization can stall training before it begins. Good initialization keeps the variance of activations and gradients roughly constant from layer to layer.

The symmetry problem

Weights must not all start equal. If every unit in a layer begins with identical weights, they receive identical gradients and update identically, so they remain identical forever, wasting the layer's capacity. Random initialization breaks this symmetry, ensuring units diverge and learn different features. Biases, by contrast, can safely start at zero.

Xavier and He

Xavier (Glorot) initialization sets weight variance to keep signal variance stable through layers with symmetric activations like tanh, scaling by the number of input and output units. He initialization adapts this for ReLU, which zeroes half its inputs, by using variance 2 divided by the number of inputs to compensate for the halving. Matching the initialization to the activation is the practical rule: He for ReLU-family, Xavier for tanh and sigmoid.

python

import numpy as np
# He initialization for a ReLU layer
W = np.random.randn(n_out, n_in) * np.sqrt(2.0 / n_in)
# Xavier initialization for tanh
W = np.random.randn(n_out, n_in) * np.sqrt(1.0 / n_in)

Interaction with normalization

Normalization layers such as batch and layer norm reduce sensitivity to initialization by rescaling activations at every step, which is part of why they made deep networks easier to train. Even so, sensible initialization still helps convergence speed and stability, and very deep or residual networks sometimes use specialized schemes, such as scaling residual branches down, to keep the early forward pass well behaved.

Practical guidance

Frameworks default to reasonable initializers, usually He or Xavier variants, so most practitioners never set them by hand. Awareness still matters when training unusually deep networks, custom layers, or architectures without normalization, where the wrong scale can prevent learning entirely. The guiding principle is simple: break symmetry with randomness, and choose a variance that preserves signal scale given the activation.