Stochastic Gradient Descent
A scalable form of gradient descent that updates parameters from small random batches of data.
Definition
Stochastic gradient descent (SGD) estimates the loss gradient from a small random subset (a mini-batch) of the training data at each step, rather than the whole dataset. This makes each update cheap and introduces useful noise that can help escape shallow minima.
The noise in each mini-batch estimate is not merely tolerated but can be beneficial, acting as an implicit regularizer that biases training toward flatter minima associated with better generalization. This is one reason very large batches, which reduce that noise, sometimes require care to match small-batch generalization.
Batch size interacts with the learning rate: larger batches reduce gradient noise and often permit larger steps, but past a point they degrade generalization unless carefully tuned. This has made large-batch training a research topic in its own right, with techniques like learning-rate warmup and scaling rules developed specifically to preserve accuracy while using the parallel hardware that big batches exploit.
One pass through the whole dataset is an epoch; each mini-batch update is an iteration.
Why batches
- Full-batch gradients are accurate but slow on large data.
- Single-example updates are noisy and inefficient on hardware.
- Mini-batches balance stability, speed, and GPU utilization.
Variants
Momentum accumulates a running average of gradients to smooth the path. Adaptive methods such as Adam scale each parameter's step by its recent gradient history, often converging faster with less tuning.
Why it matters
SGD is the practical reason deep learning scales to enormous datasets: it turns an intractable full-batch optimization into a stream of cheap, noisy updates.
Fusion connection
Mini-batch training keeps surrogate-model fitting tractable when Kronos generates large ensembles of simulated operating points for Hyperion.