Introduction to Generative AI

Discriminative vs Generative Models


You train a small network to tell handwritten 3s from handwritten 8s. It hits 99.2% accuracy on held-out data. Excellent. Then someone asks it to draw a 3.

It cannot. Not "it draws badly" — it has no mechanism at all. Its output is a single number between 0 and 1, and there is no way to run it backwards into an image.

Worse, when you inspect what it learned, it turns out the network is mostly looking at one thing: how much ink sits in the middle-left region of the canvas, because an 8 closes its left side and a 3 leaves it open. That single feature carries almost all the accuracy. The model knows nothing about the curve at the top, the stroke thickness, the slant, or the fact that digits are made of connected strokes at all. It learned the difference between two classes, which is a far smaller thing than learning what either class looks like.

This is the whole distinction in one example. A model that learns differences is discriminative. A model that learns what the data itself looks like is generative. They are answering different questions, and almost everything else follows from that.

Two different questions about the same dataDiscriminative• Learns the boundary between 3 and 8• Spends all capacity on the decision• Usually wins on plain accuracy• Cannot draw a single digitGenerative• Learns what a 3 looks like at all• Must model detail the label ignores• Gives you sampling and outlier scores• Can be turned into a classifier
The generative model learns strictly more, and pays for it in accuracy on the one question the other model was built to answer.

The two questions, stated precisely

Write xx for an input — an image, an email, a sentence — and yy for a label, such as "spam" or "digit 3".

A discriminative model learns p(y∣x)p(y \mid x): given this input, how probable is each label? It never asks how likely the input itself was. Feed it a photo of a giraffe and a digit classifier will still hand you a confident answer, because "is this input even plausible?" is not a question it can pose.

A generative model learns p(x)p(x), or p(x,y)p(x, y) when labels are present: how probable is this input, full stop? Everything else is derived from that. Because it models the input, it can be run in the other direction — sample from p(x)p(x) and you get a new input that never existed.

Discriminative models learn a boundary. Generative models learn a landscape. A boundary can sort things; only a landscape can produce them.

Picture two clouds of points on a page, one cloud per class. The discriminative model draws the line between them and throws the clouds away — it only ever needed the line. The generative model models the shape of each cloud: where its centre is, how it spreads, which directions it stretches. From the clouds you can recover the line, but from the line you cannot recover the clouds. Information flows one way only.

Working the same problem both ways

Take spam classification with a tiny vocabulary, so the arithmetic stays visible. Your training set is 100 emails: 30 spam, 70 legitimate. You track one word, free.

  • Of the 30 spam emails, 24 contain free.
  • Of the 70 legitimate emails, 7 contain free.

The discriminative route

Logistic regression looks at every email as a feature vector and fits a weight so that the training labels are predicted as well as possible. It ends up with something like: if the word free is present, push the log-odds of spam up by 2.3. Combined with a bias term of −0.85-0.85, an email containing free gets log-odds −0.85+2.3=1.45-0.85 + 2.3 = 1.45, and

p(spam∣free)=11+e−1.45=0.81p(\text{spam} \mid \text{free}) = \frac{1}{1 + e^{-1.45}} = 0.81

81% spam. The model went straight for the answer. It never formed any opinion about how often the word free appears in email generally, because it never needed one.

The generative route

Naive Bayes models each class separately, then applies Bayes' rule. From the counts:

  • p(spam)=0.30p(\text{spam}) = 0.30, p(legit)=0.70p(\text{legit}) = 0.70
  • p(free∣spam)=24/30=0.80p(\text{free} \mid \text{spam}) = 24/30 = 0.80
  • p(free∣legit)=7/70=0.10p(\text{free} \mid \text{legit}) = 7/70 = 0.10

Now combine:

p(spam∣free)=0.80×0.300.80×0.30+0.10×0.70=0.240.31=0.774p(\text{spam} \mid \text{free}) = \frac{0.80 \times 0.30}{0.80 \times 0.30 + 0.10 \times 0.70} = \frac{0.24}{0.31} = 0.774

Roughly 77%. Similar answer, arrived at through completely different bookkeeping. And notice what the generative route produced along the way that the discriminative route did not: p(free)=0.31p(\text{free}) = 0.31, the overall probability of seeing that word. That extra quantity is the whole difference.

What the extra quantity buys you

An email arrives containing words the model has essentially never seen — a message in Hungarian, say. The logistic regression still returns a number, perhaps 0.5, and offers no signal that anything is unusual. The Naive Bayes model computes p(x)p(x) and finds it astronomically low: this does not look like email from my training distribution at all. That is out-of-distribution detection, and you get it for free from a generative model. A discriminative model must have it bolted on separately.

The full comparison

PropertyDiscriminativeGenerative
Modelsp(y∣x)p(y \mid x)p(x)p(x) or p(x,y)p(x, y)
Question"Which class?""What does the data look like?"
Can create new dataNoYes
Can detect strange inputsNot nativelyYes, via low p(x)p(x)
Accuracy at a fixed task, plentiful labelsUsually higherUsually lower
Accuracy with very few labelsDegrades sharplyDegrades gently
Uses unlabelled dataNoYes — most training data can be unlabelled
Compute to trainModestSubstantially higher
Handles missing input featuresPoorlyNaturally — marginalise them out
Typical membersLogistic regression, SVM, random forest, most CNN classifiersNaive Bayes, GMM, VAE, GAN, diffusion model, large language model

The accuracy rows deserve care, because they are the source of a persistent myth.

Why discriminative usually wins on accuracy

The classic argument, due to Vapnik: when solving a problem, do not solve a harder problem as an intermediate step. Classification asks for a boundary. Modelling p(x)p(x) asks for a full description of the data — and then you throw nearly all of it away to get the boundary. Every parameter a generative model spends on faithfully capturing stroke thickness in handwritten digits is a parameter not spent on the 3-versus-8 decision.

Ng and Jordan sharpened this in 2001 with a result worth remembering. Naive Bayes (generative) and logistic regression (discriminative) on the same features behave differently as data grows:

Labelled examplesNaive BayesLogistic regressionWinner
Very few (tens)Reaches usable accuracy quicklyOverfits badlyGenerative
ModerateApproaching its ceilingStill improvingCrossover point
Many (thousands+)Plateaus below its rival — its assumptions biteKeeps improvingDiscriminative

The generative model converges faster but to a worse answer; the discriminative model converges slower but to a better one. The generative model's assumptions about the data act like a strong prior — enormously helpful when you have almost nothing to learn from, a ceiling once you have plenty.

Generative models are not worse classifiers because they are generative. They are worse because they spent their capacity answering a bigger question than the one you asked.

Turning a generative model into a classifier

Because p(x,y)p(x, y) contains everything, any generative model can be pressed into service as a classifier through Bayes' rule:

p(y∣x)=p(x∣y) p(y)∑y′p(x∣y′) p(y′)p(y \mid x) = \frac{p(x \mid y)\, p(y)}{\sum_{y'} p(x \mid y')\, p(y')}

In practice: train one generative model per class, then for a new input ask each model how likely it finds that input, and pick the class whose model is least surprised.

Python
# Classification by generative scoring, one model per classimport numpy as npdef classify(x, class_models, class_priors):    """class_models[k].log_prob(x) -> log p(x | y=k)"""    scores = [        m.log_prob(x) + np.log(prior)        for m, prior in zip(class_models, class_priors)    ]    # Softmax over log-scores gives the posterior p(y | x)    scores = np.array(scores)    posterior = np.exp(scores - scores.max())    posterior /= posterior.sum()    return int(posterior.argmax()), posterior# A useful side effect: if every model reports a very low log_prob,# the input resembles none of the training classes. Flag it rather# than forcing a label onto it.

Large language models do this constantly without anyone calling it classification. Prompting a model with "Classify the sentiment of this review as positive or negative: ..." and reading which of the two continuations it assigns higher probability is exactly Bayes-rule classification wearing a friendlier interface. It works with zero labelled examples, which no discriminative model can match, and it is usually beaten by a small fine-tuned classifier once you have a few thousand labelled examples in hand.

In practice, many hosted models do not return token probabilities, or return them only for some models; OpenAI's reasoning models, for example, do not support them. The everyday version is to ask for the label directly and use the provider's structured-output mode to restrict the answer to your fixed list of labels. The principle is the same; you just see the model's choice rather than its odds.

Choosing between them

SituationReach forReason
Fixed label set, thousands of labels, accuracy is everythingDiscriminativeBest accuracy per parameter and per training hour
You need to produce new contentGenerativeDiscriminative models structurally cannot
Labels are scarce or expensiveGenerative, or generative pre-training then a small classifier headUnlabelled data still teaches the model something
Anomaly or fraud detectionGenerative"Unlikely under p(x)p(x)" is the definition of an anomaly
Inputs arrive with fields missingGenerativeUnknown variables can be integrated out
Tight latency or edge deploymentDiscriminativeOrders of magnitude cheaper at inference
New classes appear over timeGenerativeAdd a class model; no need to retrain the boundary
You must explain a decision to a regulatorDiscriminativeFeature weights map directly to reasons

The hybrid that took over

The most important practical arrangement is not either extreme. It is generative pre-training followed by discriminative fine-tuning, and it is how virtually every strong model of the last several years was built.

  1. Pre-train generatively on a very large pile of unlabelled data. Predict masked words, predict removed noise, predict the next token. No annotator is involved. The model is forced to learn the structure of the domain, because you cannot predict a missing word well without learning grammar, facts and context.
  2. Fine-tune discriminatively on the small labelled set you actually have. Attach a classification head and train on a few thousand examples.

A rough illustration of the size of the effect: a classifier trained from scratch on 1,000 labelled sentiment examples might reach the low 80s in accuracy, while the same 1,000 examples on top of a generatively pre-trained model can reach the mid 90s. The exact numbers depend on the task, but the direction is reliable. The unlabelled corpus taught it what language is; the labels only had to teach it which end of an axis to point at.

Large language models follow the same two-stage pattern with a twist: the second stage is usually generative too. After pre-training, they are tuned on examples of instructions and good replies, then on human or automated preferences between replies. And for many classification jobs today, teams skip fine-tuning entirely and simply prompt the pre-trained model, adding a small trained classifier only when volume or accuracy demands it.

The second common hybrid is the conditional generative model — modelling p(x∣y)p(x \mid y), generation steered by a label or a prompt. Text-to-image tools, instruction-following assistants and class-conditioned image models all live here. They are generative in mechanism and controllable in the way discriminative systems are, which is why they dominate applications.

Three misconceptions worth killing

"Generative means neural." No. Naive Bayes is generative and dates to the 1960s. Gaussian mixture models are generative. Hidden Markov models are generative. The property is about what is modelled, not about how many layers are involved.

"Generative models are strictly better because they can do more." They answer a harder question with the same budget, and on a narrow task with abundant labels a gradient-boosted tree will beat them on accuracy, latency and cost simultaneously. "Can do more" is only an advantage when you need more.

"A model with high classification accuracy understands the class." The 3-versus-8 network at the start of this lesson is the counterexample. It scored 99.2% while knowing essentially one fact about ink placement. High accuracy on the boundary is compatible with near-total ignorance about the objects — which is precisely why such models break so strangely when the input distribution shifts.

What this means when you build something

Before choosing a model, write down the question you are actually asking. If it is "which of these fixed categories is this?", you want a boundary, you want it cheap, and generative machinery is overhead. If the question involves producing something, filling something in, judging whether something is normal, or coping with categories you have not enumerated yet, you need a model of the data itself and no boundary will substitute.

Then check your label budget honestly. Ten thousand clean labels point at a discriminative model. Two hundred labels and a mountain of raw unlabelled data point at generative pre-training with a small supervised layer on top — and that path will usually beat the from-scratch classifier by a wide margin.

Finally, if your system will meet inputs that nobody anticipated — and production systems always do — you need some estimate of p(x)p(x) somewhere in the stack. A classifier alone will label the giraffe as a 3, at 97% confidence, and tell you nothing is wrong.

Check your understanding

0 of 3 answered

1.A digit classifier scores 99% on its test set. You feed it a photo of a cat and it answers "7" with high confidence. Why?

2.You have 200 labelled support tickets and 500,000 unlabelled ones. Which approach is most likely to give the best classifier?

3.Why does a generative classifier such as Naive Bayes often lose to logistic regression once labelled data is plentiful?