One-Hot Encoding
Representing a categorical value as a binary vector with a single active element.
Definition
One-hot encoding represents a category as a vector the length of the vocabulary, with a 1 in the position for that category and 0 elsewhere. It lets models that expect numbers consume categorical data without implying a false ordering.
For features with high cardinality, such as identifiers with thousands of values, one-hot vectors become impractically large and sparse. Target encoding and learned embeddings are the usual alternatives, trading some interpretability for a compact, information-dense representation.
The representation implies that all categories are equidistant, which is faithful when they truly have no order but wasteful when they do; ordinal encodings preserve a meaningful ranking. For high-cardinality features, the sparsity of one-hot vectors becomes a real cost, and learned embeddings replace them by placing related categories near one another in a compact continuous space that a model can generalize over.
Trade-offs
- No spurious ordinal relationship between categories.
- Vector length grows with the number of categories.
- Sparse and memory-heavy for large vocabularies.
- Embeddings are preferred when categories are numerous.
Why it matters
One-hot encoding is the simplest correct way to feed categories into linear models and networks. Its inefficiency at scale motivated learned embeddings, which map categories into a compact continuous space.
Fusion connection
Discrete Hyperion design choices, such as a material selection, enter surrogate models as one-hot vectors, or as embeddings when the option set is large enough to warrant it.