Computing Library › Worked Examples
Worked Examples

Logistic Regression by Gradient Descent

Fit a logistic classifier to a small separable dataset with batch gradient descent and read the learned decision boundary.

Problem

Logistic regression models the probability of a binary label as a sigmoid of a linear score. Training minimizes cross-entropy loss, which is convex, so gradient descent reaches the global optimum. We fit weights on a small 1D dataset and locate the decision threshold.

Model and gradient

Prediction is p = sigmoid(w x + b). The cross-entropy gradient with respect to the weights is the average of (p - y) times the input, a remarkably clean form that follows from the sigmoid-plus-log-loss pairing.

python

import numpy as np
x=np.array([-2,-1,-0.5,0.5,1,2.]); y=np.array([0,0,0,1,1,1.])
w=b=0.0; lr=0.5
sig=lambda z:1/(1+np.exp(-z))
for _ in range(2000):
    p=sig(w*x+b)
    gw=np.mean((p-y)*x); gb=np.mean(p-y)
    w-=lr*gw; b-=lr*gb
boundary=-b/w   # where p=0.5
print('w',round(w,3),'b',round(b,3),'boundary x',round(boundary,3))

Result

Training pushes the weight positive and places the decision boundary (where p=0.5) near x=0, correctly separating the two groups. Because the data are separable the weights keep growing to sharpen the sigmoid, which is why L2 regularization is usually added to keep them finite. The learned probability is well-calibrated near the boundary and saturates toward 0 and 1 far from it.