Supervised Learning
Learning a function from input-output pairs so it predicts correct outputs for new inputs.
Definition
Supervised learning fits a model to a dataset of labeled examples, pairs of inputs x and desired outputs y. The goal is a function f(x) that predicts y accurately on inputs never seen during training.
Label quality sets a ceiling on what supervised learning can achieve, so much practical effort goes into consistent labeling and into detecting label noise. When labels are scarce, semi-supervised and self-supervised methods extend the idea by generating training signal from the structure of unlabeled data itself.
Class imbalance is a frequent complication: when one outcome is rare, a model can score high accuracy by ignoring it entirely. Remedies include resampling the data, weighting the loss toward the rare class, and evaluating with metrics that expose minority-class performance. The right choice depends on the real cost of each kind of error, which is a domain question the algorithm cannot answer on its own.
When y is a category, the task is classification; when y is a continuous quantity, it is regression. Both minimize a loss measuring the gap between prediction and label.
Common algorithms
- Linear and logistic regression
- Decision trees and random forests
- Support vector machines
- Neural networks
Why it matters
Most deployed ML is supervised because labeled data directly encodes the target behavior. The main costs are obtaining labels and ensuring the training distribution matches deployment, since a model trained on one regime degrades on another.
Fusion connection
Supervised regressors trained on simulation outputs predict quantities such as fusion power or confinement for new Hyperion configurations, standing in for slow codes during early design sweeps.