Course Content
LLMs Deep Dive
10 sections · 40 lessons
What is the difference between discriminative AI and generative AI?
What you need to know
How the engineering differs
| Discriminative AI | Generative AI | |
|---|---|---|
| Output | Label, score, rank | Text, image, audio, code |
| Correct answer | Usually one | Many acceptable answers |
| Evaluation | Accuracy, precision, recall, AUC | Rubrics, human review, LLM-as-judge, task success |
| Behaviour | Deterministic given the model | Varies with sampling |
| Main risks | Bias, drift, false positives | Hallucination, unsafe content, prompt injection, data leakage |
| Cost per call | Micro- to milliseconds, cheap | Hundreds of ms to seconds, per-token cost |
Why generative outputs are harder to trust
A fraud classifier's worst case is a wrong label, which you can measure. A generative model can produce a fluent paragraph with one wrong number in it, a policy that does not exist, or text copied from a malicious document it read. So generative systems need:
- Grounding — retrieval of the facts, with citations.
- Output checks — schema validation, moderation classifiers, numeric checks against source data.
- Input defences — treating retrieved text and user text as untrusted.
- Human review for high-risk actions.
How they combine
- Route (discriminative) — an intent classifier decides what the user wants.
- Retrieve and rank (discriminative) — a ranker picks the best documents or products.
- Generate (generative) — an LLM writes the answer from the retrieved facts.
- Check (discriminative) — a safety or policy classifier and code-based checks approve the output.
A real-life example
An e-commerce search assistant answers "Best phone under Rs 20,000 for gaming?" A discriminative intent model classifies this as a product-recommendation query. A learned ranking model scores 3,000 phones and returns the top 10 by predicted purchase likelihood and match. An LLM (generative) writes a short comparison of the top 3 using their specs from the catalogue. Finally, a checker verifies that every price and spec in the answer matches the catalogue, and a moderation classifier checks the text.
The team measures each part differently: the ranker by click-through and NDCG, the LLM by a weekly human review of 200 answers against a rubric, and the checker by how many bad answers it catches. When the LLM once wrote "under Rs 20,000" next to a phone priced Rs 21,499, the discriminative-style numeric check caught it before the shopper saw it.
Follow-up questions to expect
- "Can generative AI replace discriminative models?" — For many low-volume tasks, prompting an LLM is good enough. For high volume, strict latency or audited decisions like credit scoring, dedicated discriminative models remain better.
- "How do you evaluate a generative system?" — A fixed test set with rubrics per case, automated checks where possible (format, facts against sources), model-based judges calibrated against human ratings, and online metrics such as task success.
- "Is a recommender system generative?" — Traditionally discriminative (it scores items), though newer systems also use generative models to produce item IDs or explanations.