Course Content
LLMs Deep Dive
10 sections · 40 lessons
What are Generative and Discriminative models?
What you need to know
The probability view
- Discriminative: learn
P(y | x)— "given this message, how likely is it to be fraud?" - Generative: learn
P(x | y)andP(y), orP(x)— "what do fraud messages look like, and what do normal ones look like?" Then classify with Bayes' rule:
P(y | x) = P(x | y) × P(y) / P(x)A worked example: 2% of messages are fraud. The word "OTP" appears in 60% of fraud messages and 5% of normal ones. For a message containing "OTP":
fraud: 0.60 × 0.02 = 0.012normal: 0.05 × 0.98 = 0.049P(fraud | "OTP") = 0.012 / (0.012 + 0.049) ≈ 0.20That is how Naive Bayes, a simple generative classifier, works.
Side by side
Discriminative
- Models P(label given input)
- Learns only the decision boundary
- Needs labelled data
- Usually best accuracy for classification
Generative
- Models the data itself, P(input) or P(input, label)
- Can create new samples
- Can learn from unlabelled data
- Harder to train; more to learn
Where LLMs fit
An LLM is generative: it models P(text) as a product of next-token probabilities. But you can use it as a classifier by asking a question and reading the probability of each label token, which is why prompting can replace a trained classifier. Conversely, BERT is pretrained with a generative-style objective (recovering masked tokens) and then fine-tuned as a discriminative classifier.
An everyday framing
A real-life example
A bank needs to flag suspicious UPI messages. Option one: a discriminative gradient-boosted model trained on 500,000 labelled messages. It runs in 2 ms and gives well-calibrated fraud probabilities. Option two: a generative LLM asked "Is this a phishing message? Answer yes or no." It needs no labelled data, and catches new scam wording, but costs far more per message and takes hundreds of milliseconds.
The bank uses both: the discriminative model scores every message, and the LLM is called only on the 3% the model is unsure about, adding an explanation for the fraud analyst. When the fraud team lacked labels for a new scam type, they used the LLM (generatively) to write 2,000 synthetic examples of that scam, reviewed them, and used them to retrain the discriminative model.
Follow-up questions to expect
- "Is logistic regression generative or discriminative?" — Discriminative; it models P(y given x) directly. Its generative counterpart is Naive Bayes.
- "Why can generative models handle missing features?" — They model the joint distribution, so they can sum over the missing values; a discriminative model expects every input.
- "Is a GAN's discriminator discriminative?" — Yes: it classifies real versus generated, while the generator is the generative part.