Gradient Boosting
An ensemble method that builds trees sequentially, each correcting the residual errors of the ensemble so far.
Definition
Gradient boosting builds an ensemble of decision trees one at a time. Each new tree is fit to the gradient of the loss with respect to the current ensemble's predictions, so it focuses on the examples the ensemble handles worst.
The learning rate and the number of trees trade off directly: a smaller rate needs more trees but often generalizes better. Early stopping on a validation set is the standard way to choose the number of trees automatically and avoid overfitting.
Modern implementations add engineering that matters as much as the algorithm: histogram-based splitting for speed, native handling of missing values, and built-in regularization on tree structure. These refinements are why gradient-boosted trees remain the leading method for tabular data even as deep learning dominates images and text, and why they are a strong default when the data arrives as rows and columns.
The final prediction is a weighted sum of all trees. A learning rate shrinks each tree's contribution to avoid overshooting.
Key knobs
- Number of trees and learning rate (trade off jointly).
- Tree depth, controlling interaction complexity.
- Subsampling of rows and columns for regularization.
Why it matters
Implementations such as XGBoost and LightGBM are among the most accurate methods for structured tabular data and dominate many prediction competitions. They require more tuning and are more prone to overfitting than random forests if pushed too hard.
Fusion connection
Gradient-boosted models serve as high-accuracy surrogates for tabular Hyperion outputs when squeezing out predictive accuracy justifies the extra tuning.