Machine Learning Essentials

Logistic Regression


You have 2,000 loan applications, each with an income figure and a column saying whether the borrower defaulted. You want to predict default for new applicants. You already know how to fit a line, so try that: code the outcome as 0 for repaid and 1 for defaulted, and run linear regression.

It runs. It produces a line. And then you look at the predictions.

Text
income      prediction£18,000        1.34£25,000        1.02£40,000        0.61£70,000        0.14£95,000       -0.19£140,000      -0.62

A predicted probability of 1.34 is not a probability. Neither is −0.62. You could clip them to the range [0, 1], but that only papers over the problem — and there is a second, worse one.

Add a single applicant earning £900,000 who repaid. In a probability model this should barely matter: one more low-risk person among many. In least squares it is a point far out along the x-axis with enormous leverage, and it tilts the entire line. Applicants around £40,000, who were previously predicted at 0.61, now come out at 0.44 — their predicted risk changed substantially because of one person whose situation has nothing to do with theirs.

Both failures come from the same root. Linear regression models a quantity that runs from −∞-\infty to +∞+\infty, and you are asking it to model something confined to [0, 1]. The shape is wrong. Not the coefficients — the shape.

From a linear score to a decisionWeighted sumof featuresSigmoidsquashesit to 0 to 1Read it as aprobabilityCompareagainst athresholdPredicted classOnly the last arrow uses the threshold, so the same fitted model can be retuned without refitting.
The model outputs a probability; the 0.5 cut-off is a business choice bolted on afterwards.

What shape do we actually want?

Think about how default probability should behave as income rises from £5,000 to £5,000,000.

At the very low end, probability should be high and should flatten — the difference between earning £5,000 and £8,000 barely changes an already dire situation. At the very high end it should be low and flatten again — £900,000 and £1.2m are both simply "rich". In the middle, around wherever the decision actually tips, probability should change fast: the difference between £30,000 and £45,000 is where real discrimination lives.

That description is an S-curve: flat, steep, flat. Bounded at both ends, steepest in the middle. And there is a standard function with exactly that shape.

The sigmoid, and where it comes from

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

zz−6−3−10136
σ(z)\sigma(z)0.0020.0470.2690.5000.7310.9530.998

It takes any real number and returns something strictly between 0 and 1, never reaching either. Very negative inputs map to near-certainty of "no", very positive to near-certainty of "yes", and zero maps to exactly 0.5.

So the model becomes: compute a linear combination as before, then squash it.

z=β0+β1x1+⋯+βpxp,P(y=1∣x)=σ(z)=11+e−zz = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p, \qquad P(y = 1 \mid \mathbf{x}) = \sigma(z) = \frac{1}{1 + e^{-z}}

This looks like an arbitrary patch — take a broken model and bend its output into range. It is not. Rearranging shows something cleaner. Start from p=σ(z)p = \sigma(z) and solve for zz:

z=ln⁡(p1−p)z = \ln\left(\frac{p}{1-p}\right)

The quantity p/(1−p)p/(1-p) is the odds: the ratio of "happens" to "doesn't happen". A probability of 0.75 is odds of 3, usually written 3-to-1. Its logarithm is the log-odds, or logit.

So the model is not "linear regression with a squashing function bolted on". It is exactly linear regression — performed on the log-odds instead of the probability:

ln⁡(P(y=1)P(y=0))=β0+β1x1+⋯+βpxp\ln\left(\frac{P(y=1)}{P(y=0)}\right) = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p

Probability is bounded and behaves awkwardly. Log-odds run from −∞-\infty to +∞+\infty and behave linearly. Logistic regression is the observation that if you change what you are modelling, the straight line works fine.

Reading coefficients in odds

This gives coefficients a clean interpretation. Since zz is log-odds and it increases by βj\beta_j per unit of xjx_j:

  • A one-unit increase in xjx_j adds βj\beta_j to the log-odds.
  • Equivalently, it multiplies the odds by eβje^{\beta_j}.

Suppose a churn model gives these coefficients:

Featureβ\betaeβe^{\beta}Plain English
support_tickets+0.621.86Each extra ticket multiplies the odds of churn by 1.86 — an 86% increase in odds
tenure_years−0.410.66Each extra year multiplies churn odds by 0.66 — a 34% reduction
has_contract−1.200.30Being on contract cuts the odds of churn to 30% of otherwise
monthly_spend+0.0041.004Each extra £1 raises the odds by 0.4%

The critical caution: a change in odds is not a change in probability. Doubling the odds when the base probability is 0.02 takes it to about 0.04. Doubling the odds when the base is 0.60 takes it to 0.75. The same coefficient produces a different probability change depending on where you start, which is precisely what the S-curve encodes — steep in the middle, flat at the ends.

The decision boundary is still a straight line

Predict class 1 when σ(z)≥0.5\sigma(z) \geq 0.5. Since σ(z)=0.5\sigma(z) = 0.5 exactly when z=0z = 0, the boundary is:

β0+β1x1+β2x2=0\beta_0 + \beta_1 x_1 + \beta_2 x_2 = 0

which is a straight line in two dimensions, a plane in three, a hyperplane in general. The sigmoid changed how confidence is expressed, not the geometry of the decision. Logistic regression is a linear classifier, and it cannot separate classes that are not linearly separable — the classic example being two concentric rings, which no straight line divides.

The fix is the same as for regression: engineer nonlinear features. Adding x12+x22x_1^2 + x_2^2 as a feature lets a linear boundary in the expanded space become a circular boundary in the original one.

Why not just use squared error?

Having chosen the model, you need a loss. The obvious candidate — squared error between predicted probability and the 0/1 label — fails for two reasons.

The first is mathematical: composing the sigmoid with squared error produces a non-convex loss surface, with local minima where gradient descent gets stuck. There is no guarantee of finding the best coefficients.

The second is about incentives, and it matters more. Under squared error, a confident wrong prediction is barely punished. Predict 0.99 when the truth is 0 and the squared error is 0.98 — bounded, and not much worse than a hedged guess. A model that is confidently, catastrophically wrong ought to be punished far more than one that shrugs.

Cross-entropy (log loss) does this:

J=−1n∑i=1n[yiln⁡(p^i)+(1−yi)ln⁡(1−p^i)]J = -\frac{1}{n}\sum_{i=1}^{n}\left[y_i \ln(\hat{p}_i) + (1 - y_i)\ln(1 - \hat{p}_i)\right]

Only one term is ever active. When y=1y = 1 the loss is −ln⁡(p^)-\ln(\hat{p}); when y=0y = 0 it is −ln⁡(1−p^)-\ln(1 - \hat{p}).

True labelPredicted prob.Squared errorLog loss
10.990.00010.01
10.700.090.36
10.500.250.69
10.100.812.30
10.010.984.61
10.0010.9986.91

Squared error runs out of road — it can never exceed 1. Log loss goes to infinity as confidence in the wrong answer approaches certainty. That unbounded punishment is what forces the model to output honest probabilities rather than bravado. It also makes the loss convex, so there is a single global minimum.

There is no closed form

Unlike least squares, no algebraic formula gives the coefficients. They are found iteratively. Scikit-learn's solver parameter chooses the method:

SolverHandlesUse when
lbfgs (default)L2, no penaltySmall to medium data — a good default
liblinearL1, L2Small data, high dimensions, sparse text features
sagaL1, L2, elastic netLarge data; the only solver supporting elastic net
newton-choleskyL2Many rows, few features

Since scikit-learn 1.8 you choose the penalty with l1_ratio rather than the old penalty= argument, which is deprecated: l1_ratio=0 (the default) is L2, l1_ratio=1 is L1, a value in between is elastic net, and C=np.inf switches the penalty off.

If you see ConvergenceWarning, the optimiser ran out of iterations. Raise max_iter, but first check that your features are scaled — unscaled features are the usual cause, because they distort the loss surface into a long narrow valley.

The threshold is a knob, not a constant

A 0.5 cutoff is a default with no special authority. The model outputs a probability; converting it to a decision is a separate step with its own economics.

Fraud detection makes this vivid. Suppose 0.3% of transactions are fraudulent, a missed fraud costs £220 on average, and a false alarm costs £4 in review time.

ThresholdFlaggedFraud caughtFalse alarmsMissed costReview costTotal
0.901209822£44,440£88£44,528
0.50410212198£19,360£792£20,152
0.20980268712£7,040£2,848£9,888
0.082,1002891,811£2,420£7,244£9,664
0.034,9002974,603£660£18,412£19,072

The cost-optimal threshold is around 0.08, not 0.5, and using the default would have cost this business roughly twice as much. Nothing about the model changed — only the number you compare its output to.

Python
import numpy as npprobs = model.predict_proba(X_val)[:, 1]     # column 1 = P(class 1)best = min(    ((220 * ((y_val == 1) & (probs < t)).sum()      + 4 * ((y_val == 0) & (probs >= t)).sum(), t)     for t in np.arange(0.01, 1.0, 0.01)))print("cheapest threshold:", best[1], "cost:", best[0])

Tune the threshold on validation data, never on the test set, and never on the training set — the model is overconfident on data it has memorised, which pushes the chosen threshold in the wrong direction.

Regularisation, and the parameter that reads backwards

Scikit-learn regularises logistic regression by default, which is sensible but catches people out because the parameter is inverted. C is the inverse of regularisation strength:

CRegularisationEffect
0.001Very strongCoefficients crushed towards zero; likely underfits
1.0 (default)ModerateReasonable starting point
1000Almost noneClose to unpenalised; may overfit badly

Because there is a penalty on coefficient size, feature scaling is not optional. A feature measured in pounds and one measured in years get penalised on completely different scales, and the model will effectively ignore whichever happens to have small numbers.

Perfect separation

One failure specific to logistic regression: if some feature separates the classes perfectly — say every applicant with prior_default = 1 defaulted — then the likelihood is maximised by driving that coefficient to infinity. Unpenalised fitting will not converge, and coefficients will grow without limit until the iteration cap stops them. The symptoms are a convergence warning plus a coefficient like 47.3.

Regularisation fixes this automatically, which is a good reason to leave it on. But a perfectly separating feature is also a strong hint of leakage, so investigate before accepting it.

More than two classes

Two strategies, and scikit-learn's default is now the second.

One-vs-rest trains one binary model per class ("is it a cat or not?", "is it a dog or not?") and picks the highest score. Simple, but the probabilities come from independent models and do not sum to 1 without rescaling.

Multinomial (softmax) trains one model producing all class scores jointly and normalises them:

P(y=k∣x)=ezk∑j=1KezjP(y = k \mid \mathbf{x}) = \frac{e^{z_k}}{\sum_{j=1}^{K} e^{z_j}}

This is a proper probability distribution by construction and generally gives better-calibrated results.

A full example

Python
import numpy as npimport pandas as pdfrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.metrics import (classification_report, roc_auc_score,                             brier_score_loss)df = pd.read_csv("churn.csv")X = df[["tenure_years", "monthly_spend", "support_tickets", "has_contract"]]y = df["churned"]X_train, X_test, y_train, y_test = train_test_split(    X, y, test_size=0.2, stratify=y, random_state=42)model = make_pipeline(    StandardScaler(),    LogisticRegression(C=1.0, class_weight="balanced", max_iter=1000),).fit(X_train, y_train)probs = model.predict_proba(X_test)[:, 1]print(classification_report(y_test, (probs >= 0.5).astype(int), digits=3))print("ROC AUC:", round(roc_auc_score(y_test, probs), 3))print("Brier  :", round(brier_score_loss(y_test, probs), 4))  # lower is better# coefficients as odds ratios, on standardised featurescoefs = pd.Series(model[-1].coef_[0], index=X.columns)print((np.exp(coefs)).sort_values(ascending=False).rename("odds ratio per SD"))

class_weight="balanced" tells the model to weight the rare class more heavily, which matters when positives are scarce; without it, a model facing 3% positives can minimise loss quite effectively by predicting "no" for nearly everything.

Calibration is the underrated advantage

Because logistic regression is fitted by maximising the likelihood of the observed labels, its probabilities tend to be well calibrated: among the cases it assigns 0.30, roughly 30% really are positive. That is not true of many stronger classifiers — an unmodified random forest or SVM will happily output a "probability" of 0.9 for a group where the true rate is 0.6.

When the output feeds a downstream calculation — expected loss, a pricing decision, a ranking with a cost threshold — calibration matters as much as ranking quality, and this is where a simple model often beats a more accurate one.

What this means when you build something

Fit logistic regression first on any classification problem, before anything more sophisticated. It trains in under a second, gives you a baseline that a complex model has to beat by enough to justify its cost, and produces coefficients you can put in front of a colleague and defend.

When you fit it, do three things that people routinely skip. Scale the features, because the default regularisation makes unscaled inputs meaningless. Use predict_proba rather than predict, so you keep the probability and can choose your own threshold instead of accepting 0.5. And convert the coefficients to odds ratios with np.exp before showing anyone, because "multiplies the odds by 1.86" is a sentence people understand and "coefficient of 0.62" is not.

Where it will let you down is when the true boundary is genuinely curved, or when interactions between features matter and you have not built them by hand. The signal is a model that plateaus at mediocre performance no matter how you tune it while a tree-based model immediately does better. That is not a reason to skip fitting it — it is precisely how you find out that your problem has structure a straight line cannot see.