Computing Library › Neural Architectures
Neural Architectures

Contrastive Learning

Contrastive learning trains representations by pulling similar examples together and pushing dissimilar ones apart, often without labels.

Learning by comparison

Contrastive learning shapes an embedding space through comparison rather than direct labels. Given an anchor example, the model is trained so that a matching positive example lands close to it and unrelated negative examples land far away. The model never needs to name what an example is; it only needs to judge what is similar. This makes contrastive methods a powerful form of self-supervised representation learning.

Creating positive pairs

In self-supervised vision, positive pairs are two augmented views of the same image, for example a crop and a color-jittered version. The model learns that these should map to nearby vectors despite surface differences, forcing it to capture the underlying content. Negatives are simply other images in the batch. This turns unlabeled data into a rich training signal built from the structure of the data itself.

The InfoNCE loss

The standard objective is InfoNCE, a softmax over similarities. For an anchor, it maximizes the similarity to its positive relative to the similarities to all negatives, effectively a classification of which item is the true match. A temperature parameter scales the similarities and controls how sharply the loss distinguishes hard negatives from easy ones. Larger batches supply more negatives and generally improve the learned representations.

python

# InfoNCE for an anchor a and positive p among negatives n_i
# sim(x,y) = cosine(x, y) / temperature
# loss = -log( exp(sim(a,p)) / (exp(sim(a,p)) + sum_i exp(sim(a,n_i))) )

Frameworks and refinements

SimCLR relies on large batches for negatives; MoCo maintains a queue of past embeddings so negatives need not all fit in one batch. Some methods, such as BYOL, learn good representations without explicit negatives by using a slowly updated target network, sidestepping the need to collect many negatives. Supervised contrastive learning uses labels to define positives as same-class examples, combining contrast with supervision.

Cross-modal and applications

CLIP applies contrastive learning across modalities, training image and text encoders so that an image and its caption embed nearby while mismatched pairs embed far apart. The result is a shared image-text space that enables zero-shot classification and semantic search. More broadly, contrastive pretraining produces transferable features for downstream tasks and can align representations across sensors or data types, useful wherever labels are scarce but similarity is easy to define.