Gradient Descent
An optimization method that iteratively steps parameters in the direction that most reduces a loss.
Definition
Gradient descent minimizes a function by repeatedly moving its parameters opposite to the gradient, the direction of steepest increase. Each step is scaled by a learning rate that controls how far the parameters move.
The learning rate is the single most important hyperparameter to get right, and schedules that decrease it over training, such as warmup followed by decay, are standard practice. Second-order methods use curvature information for faster convergence but are usually too costly for the largest models.
Non-convex loss surfaces mean the method finds a local, not global, minimum, yet in large networks many local minima yield similar, good performance, and the practical worry is more about saddle points and poorly conditioned regions that slow progress. Adaptive optimizers and learning-rate schedules exist largely to navigate this landscape efficiently rather than to guarantee a global optimum that is neither needed nor attainable.
For a loss L and parameters w, the update is w := w - eta * dL/dw, where eta is the learning rate. Iterating this drives the parameters toward a local minimum.
Practical concerns
- Too large a learning rate overshoots and diverges.
- Too small a rate converges slowly.
- Momentum and adaptive methods (Adam) accelerate convergence.
- Non-convex losses have many local minima and saddle points.
Why it matters
Gradient descent and its variants are the standard training engine for nearly all modern machine learning, from linear regression to large neural networks.
Fusion connection
Gradient-based optimization over differentiable surrogates lets Kronos engineers tune many Hyperion parameters at once toward a performance target, far faster than one-at-a-time scans.