Logistic Regression by Gradient Descent
Fit a logistic classifier to a small separable dataset with batch gradient descent and read the learned decision boundary.
Problem
Logistic regression models the probability of a binary label as a sigmoid of a linear score. Training minimizes cross-entropy loss, which is convex, so gradient descent reaches the global optimum. We fit weights on a small 1D dataset and locate the decision threshold.
Model and gradient
Prediction is p = sigmoid(w x + b). The cross-entropy gradient with respect to the weights is the average of (p - y) times the input, a remarkably clean form that follows from the sigmoid-plus-log-loss pairing.
python
import numpy as np
x=np.array([-2,-1,-0.5,0.5,1,2.]); y=np.array([0,0,0,1,1,1.])
w=b=0.0; lr=0.5
sig=lambda z:1/(1+np.exp(-z))
for _ in range(2000):
p=sig(w*x+b)
gw=np.mean((p-y)*x); gb=np.mean(p-y)
w-=lr*gw; b-=lr*gb
boundary=-b/w # where p=0.5
print('w',round(w,3),'b',round(b,3),'boundary x',round(boundary,3))Result
Training pushes the weight positive and places the decision boundary (where p=0.5) near x=0, correctly separating the two groups. Because the data are separable the weights keep growing to sharpen the sigmoid, which is why L2 regularization is usually added to keep them finite. The learned probability is well-calibrated near the boundary and saturates toward 0 and 1 far from it.
- Cross-entropy loss is convex in the weights, so any local minimum is global.
- On perfectly separable data unregularized weights diverge; regularization or early stopping controls this.
- Logistic regression is a strong, interpretable baseline Kronos uses before reaching for larger models on diagnostic classification.