Machine learning vs deep learning vs generative AI: which to use

JR

Jai Rao

August 22, 202619 min read

Two of the three are nested and one is a mode of use. How to choose by data type, labelled rows, latency and audit needs — and why boosting still beats a neural net on tables.


The three terms are usually taught as a ladder of sophistication: machine learning is the old thing, deep learning is the serious thing, generative AI is the current thing. That ordering smuggles in a claim — that the newest option is the strongest and the only real question is how ambitious you are willing to be — and the claim is false often enough to be expensive. Two of the three are nested inside one another, the third is a mode of using either, and on a large class of real problems the winner is ordinary deterministic code with no model in it at all.

What follows separates them by mechanism rather than definition, then converts that into a choice you can defend when someone asks why you did not use a neural network. The answer, which the rest of the post earns: your data type and the number of labelled rows you can actually put your hands on decide this for you, and they usually point somewhere less fashionable than you were hoping.

Nesting, not ranking: how the four options relate

Start with the option nobody puts in the comparison, because it is the baseline every model has to beat: rule-based software. A human writes the mapping from input to output. Tax brackets, shipping-fee tiers, a validation check, a routing table. It is deterministic, correct by construction on every case the author thought about, runs in microseconds, and is completely blind outside its rules.

Machine learning replaces the act of writing that mapping. You supply examples of input paired with desired output, choose a family of functions, and let an optimiser search that family for parameters minimising a loss. The defining property is not intelligence or scale — it is that the mapping is fitted rather than written. Logistic regression, decision trees, boosted ensembles, support vector machines, k-nearest neighbours and neural networks are all machine learning in exactly the same sense.

Deep learning is a strict subset of that. It is machine learning where the function is a stack of parameterised layers, deep enough that the intermediate layers end up holding a learned representation of the input, and where every layer is trained together by propagating gradients back through the whole stack. Nothing in deep learning sits outside machine learning. "Deep" describes composition depth, not capability tier.

Generative AI is not a third sibling at the same level. It describes what the model produces. A discriminative model maps an input to a label or a number: this transaction is fraudulent, this image contains a cat. A generative model instead captures enough of the distribution the data came from that you can draw new items out of it. That is a property of the objective, not the architecture — Gaussian mixture models and n-gram language models are generative and neither is deep, while diffusion image models and transformer language models are generative and both are. What the market calls generative AI is one corner of the space: very large deep generative models over text, images and audio.

So: one set inside another, with a third cutting across both — and the whole diagram sitting beside a box marked "just write the code" that wins more often than the diagram suggests.

Classical ML means a human chooses what the model sees

The most useful thing to understand about classical machine learning is that a person decides what the model is allowed to look at. Your raw data is a JSON record, a database row, a log line, and a model cannot consume any of those. Someone has to turn it into a fixed-width row of numbers, and every number in that row is a hypothesis about what matters.

Here is that step in full, for a late-delivery prediction problem. Look at what the function keeps and what it silently discards.

Text
from datetime import datetime, dateorder = {"placed_at": "2026-03-11T19:42:00", "customer_since": "2025-08-02",         "items": [{"sku": "TEA-250", "qty": 2, "price": 349.0},                   {"sku": "MUG-01", "qty": 1, "price": 199.0}],         "ship_to": {"pincode": "560001", "distance_km": 12.4}}def features(o):    ts = datetime.fromisoformat(o["placed_at"])    qty = sum(i["qty"] for i in o["items"])    value = sum(i["qty"] * i["price"] for i in o["items"])    return {        "hour": ts.hour,        "is_evening_peak": int(17 <= ts.hour <= 20),        "weekday": ts.weekday(),        "n_items": len(o["items"]),        "total_qty": qty,        "order_value": value,        "mean_unit_price": value / max(qty, 1),        "tenure_days": (ts.date() - date.fromisoformat(o["customer_since"])).days,        "distance_km": o["ship_to"]["distance_km"],    }print(features(order))# {'hour': 19, 'is_evening_peak': 1, 'weekday': 2, 'n_items': 2, 'total_qty': 3,#  'order_value': 897.0, 'mean_unit_price': 299.0, 'tenure_days': 221, 'distance_km': 12.4}

Nine numbers came out. The SKU strings, the pincode, the nesting, the exact timestamp — gone. And is_evening_peak is not a measurement, it is a belief: someone decided that evening traffic matters and that 17:00 to 20:00 is where the effect lives. That is domain knowledge entering the model through the only door it has.

The consequence is the underrated part. Your accuracy ceiling is set by the feature set, not the model. Swapping logistic regression for a boosted ensemble moves you a few points inside that ceiling; a feature you failed to compute is a wall. If lateness is really driven by one courier partner's shift handover and no column encodes courier or handover time, no amount of tuning finds it — the information is not in the room. Which is why experienced practitioners spend their time on the feature pipeline and very little on model selection.

Why axis-aligned splits keep winning on tables

A decision tree only ever asks one kind of question: is column j above threshold t. Gradient boosting builds a few hundred small trees in sequence, each one fitted to the errors the ensemble has made so far, and adds them up. That is the whole mechanism, and its match to tabular data is not an accident.

Table columns are heterogeneous and share no geometry. "Distance in kilometres" and "region code 7" do not live in the same space, so a linear combination of them means nothing — yet a linear combination of every column is precisely the first operation a dense neural layer performs. The network must learn its way out of a coordinate system that was never meaningful, by gradient descent, from your data. A tree never enters that trap: it looks at one column at a time, in that column's own units, so every boundary it draws is axis-aligned rather than a diagonal slice through invented coordinates. Four more properties fall out of the same design:

  • Monotone transforms are free. A split on "distance above 10" survives rescaling, log-transforming or standardising. No normalisation step to get wrong, and none to reproduce exactly at serving time.
  • Thresholds match how business processes behave. An SLA cutoff, a pricing tier, a free-shipping minimum: a tree represents a step with one split, while a smooth network approximates it with many units and never quite gets the corner.
  • Useless columns cost almost nothing. A split on a noise column never gets chosen, because it does not reduce the loss. A dense layer has to learn to drive those weights to zero instead, and it does that imperfectly.
  • Missing values are a branch, not a crisis. Modern implementations send them down a learned side of the split — no imputation, no missingness indicator to invent.

The claim is easy to test. This builds a table that looks like a business table — wildly different scales, integer category codes, threshold effects, ten columns of pure noise — and fits both a boosted ensemble and a small network to it.

Text
import numpy as np, timefrom sklearn.ensemble import HistGradientBoostingClassifierfrom sklearn.neural_network import MLPClassifierfrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerdef make_table(n, seed=0):    rng = np.random.default_rng(seed)    tenure_days = rng.integers(0, 3000, n)      # 0..3000    order_value = rng.lognormal(4.0, 1.1, n)    # long tail, ~5..2000    distance_km = rng.exponential(8, n)         # 0..60    hour        = rng.integers(0, 24, n)        # 0..23    region      = rng.integers(0, 12, n)        # arbitrary integer codes    junk        = rng.normal(0, 1, (n, 10))     # ten columns of nothing    score = (1.2 * (distance_km > 10)             + 1.0 * ((hour >= 17) & (hour <= 20))             + 0.9 * (tenure_days < 400)             + 1.1 * ((order_value > 100) & (region % 3 == 0)))    z = score + rng.normal(0, 0.5, n)    y = (z > np.median(z)).astype(int)          # balanced by construction    return np.column_stack([tenure_days, order_value, distance_km, hour, region, junk]), yX, y = make_table(5000)Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)for name, model in [    ("gradient boosting", HistGradientBoostingClassifier(random_state=0)),    ("MLP (64, 64)", make_pipeline(StandardScaler(),                     MLPClassifier((64, 64), max_iter=800, random_state=0))),]:    t = time.perf_counter(); model.fit(Xtr, ytr); fit = time.perf_counter() - t    print(f"{name:18s} acc={model.score(Xte, yte):.3f}  fit={fit:5.2f}s")

On a laptop CPU, gradient boosting lands around 0.84 in roughly a third of a second; the network lands near 0.67 and takes several seconds. Plain logistic regression on the same columns reaches about 0.70, so the network's extra capacity bought nothing over a linear model. Being fair to it does not rescue it: quantile-transforming the numeric columns, one-hot encoding the category and going up to two hidden layers of 256 units still lands around 0.70. That is an afternoon of preprocessing and hyperparameters spent drawing level with a five-line linear baseline the ensemble beat by fourteen points before the coffee was ready.

Data volume is the other half of the story, and it is the half that gets skipped. Reuse that generator at different row counts and refit both models:

Text
for n in [500, 2000, 20000, 100000]:    X, y = make_table(n)    Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.25, random_state=0, stratify=y)    g = HistGradientBoostingClassifier(random_state=0).fit(Xtr, ytr)    m = make_pipeline(StandardScaler(), MLPClassifier((128, 64), max_iter=400,        early_stopping=True, random_state=0)).fit(Xtr, ytr)    print(f"n={n:6d}  gbm={g.score(Xte, yte):.3f}  mlp={m.score(Xte, yte):.3f}")

At 500 rows the ensemble is already around 0.83 — within four points of where it lands with two hundred times more data. The network starts near 0.73, is still around 0.71 at 5,000 rows, and needs 100,000 rows to reach roughly 0.82, where it is still losing. That is the practical meaning of "needs less data": each split is estimated from many rows at once, so a tree is cheap in examples per unit of structure it captures. A network estimating its own representation from scratch has far more to pin down and only your labels to pin it down with. Push it to 200,000 rows and a fully trained network gets within about a point of the ensemble — a minute of fitting against the ensemble's half second, on forty times the data, and still not winning.

Inspectability closes the case. Run sklearn.inspection.permutation_importance on the fitted ensemble and it ranks distance, hour, tenure, order value and region at the top with all ten noise columns at roughly zero — it recovers the actual generating rules, in order. A single prediction decomposes into a readable path: distance above 10, hour in the evening block, tenure under 400 days. Ask a network the same question and you get an attribution method, a heatmap, and an argument about whether the attribution method can be trusted.

Deep learning trades feature engineering for data and compute

Now take a problem where the feature-engineering step collapses: classify a photograph. There is no sensible list of numbers a person can write down, because the information lives in relationships among thousands of pixels and there are vastly more useful relationships than anyone could enumerate. The same holds for a raw waveform and for a sequence of words.

Deep learning's actual shift is to stop supplying features and supply raw signal plus an architecture whose structure suits the signal. Weight sharing and locality for images, because a nose is a nose wherever it appears. Attention over positions for sequences, because meaning depends on which earlier tokens matter. Gradient descent then tunes the intermediate layers until the final layer's job is nearly trivial. Those intermediate activations are the features — nobody named or specified them, and in a vision model the early ones reliably come out as edge and colour-blob detectors, because that is what minimising the loss demands.

The clearest way to feel the difference is to hand-craft features for a problem where they are hopeless and compare against handing over the raw signal. These are 8x8 handwritten digits; the hand-built set is nine features a reasonable person would think of — ink per quadrant, total ink, left-right and top-bottom asymmetry, centre of mass.

Text
import numpy as npfrom sklearn.datasets import load_digitsfrom sklearn.linear_model import LogisticRegressionfrom sklearn.neural_network import MLPClassifierfrom sklearn.model_selection import train_test_splitfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerd = load_digits(); imgs = d.images / 16.0        # 1797 images, 8x8, scaled to 0..1def handmade(img):                               # nine features a person thought of    quads = [img[:4, :4].sum(), img[:4, 4:].sum(), img[4:, :4].sum(), img[4:, 4:].sum()]    rows, cols = img.sum(1), img.sum(0)    return quads + [img.sum(),                    np.abs(img - img[:, ::-1]).sum(),   # left-right asymmetry                    np.abs(img - img[::-1, :]).sum(),   # top-bottom asymmetry                    (rows * np.arange(8)).sum() / (img.sum() + 1e-9),                    (cols * np.arange(8)).sum() / (img.sum() + 1e-9)]Xhand = np.array([handmade(i) for i in imgs])    # 9 engineered featuresXraw  = imgs.reshape(len(imgs), -1)              # 64 raw pixelsfor label, X, est in [("hand-built + logreg", Xhand, LogisticRegression(max_iter=3000)),                      ("hand-built + MLP",    Xhand, MLPClassifier((128, 64), max_iter=1500, random_state=0)),                      ("raw pixels + logreg", Xraw,  LogisticRegression(max_iter=3000)),                      ("raw pixels + MLP",    Xraw,  MLPClassifier((128, 64), max_iter=1500, random_state=0))]:    a, b, c, e = train_test_split(X, d.target, test_size=0.25, random_state=0, stratify=d.target)    print(f"{label:22s} acc={make_pipeline(StandardScaler(), est).fit(a, c).score(b, e):.3f}")

The hand-built features cap out around 0.73 with a linear model and 0.81 with a network. Raw pixels reach about 0.97 — with either model. The lesson is not "the network won": it barely helped, because 8x8 digits are close to linearly separable once you keep all the pixels. The lesson is that the ceiling was set by the inputs. Nine thoughtful numbers threw away information no downstream model could recover, and the fix was to stop summarising and let the model decide what to look at. Scale up to real photographs and the linear model on raw pixels does collapse — that is where a convolutional stack earns its cost, because now the representation itself has to be learned.

The bill arrives in four parts. Labelled data, because learning a representation from scratch means fitting far more parameters with nothing but labels to constrain them — which is why starting from a pretrained backbone with a small head on top is now the default. Compute, in GPU hours to train and accelerator cost per request to serve. Opacity, since there is no feature list to hand an auditor. And fragility: a model that learned its own features can latch onto a correlate of the label — the scanner, the watermark, the background — and you will not spot it in a column listing, because there is no column listing.

The output stops being a label

The generative shift is a change of target. A discriminative model estimates the probability of a label given an input and returns one element of a fixed set. A generative model learns enough about the distribution the data came from to produce a new item from it — a next token conditioned on everything so far, or an image obtained by repeatedly denoising a noise field toward something that looks like the training data. Prediction becomes production.

Four consequences matter when you are choosing:

  • The output is content, not a decision. There is no single right answer sitting beside it, so checking correctness stops being an accuracy computation and becomes a judgement about whether the artefact is well-formed and acceptable.
  • Identical input can give different output, by design. Sampling is the mechanism, not a defect.
  • The output space is unbounded, so anything downstream has to defend itself: schema validation, allowed-value checks, a human in the path for consequential actions.
  • Cost and latency scale with how much you generate, not just with what you send. A long answer costs more and takes longer, every single call.

Generative is a different job, not a higher rung. If your product needs to return a number, a class or a decision, and you route it through a generative model, you have taken a discriminative problem and solved it in the least verifiable way available to you.

The constraints that actually decide it

Held against the things that genuinely constrain a system, the four separate cleanly.

ConstraintRule-based codeClassical MLDeep learningDeep generative
Data it suitsEnumerable, stable rulesTables: named, heterogeneous columnsRaw perceptual signal: pixels, audio, textOpen-ended content over those signals
Labelled examplesNoneHundreds to tens of thousandsTens of thousands up, or a pretrained base to fine-tuneNone of your own if you rent a model; enormous corpora to train one
InterpretabilityTotal: read the codeGood: named features, importances, readable split pathsPoor: attribution methods, not explanationsPoor, and non-deterministic on top
Training costNoneSeconds to minutes on a CPUGPU hours to weeks; fine-tuning much lessNot yours, unless you are training foundation models
Typical latencyMicrosecondsWell under a millisecond per rowMilliseconds to tens of milliseconds with an acceleratorHundreds of milliseconds to tens of seconds, growing with output length

Nothing in that table says "more advanced". Each column is better than its neighbours at something specific and worse at everything else.

Picking one from what you already have

Work through these in order and stop at the first that answers. The order runs from cheapest and most verifiable to most expensive and least verifiable — the opposite of the order most teams use.

  1. Can you write the rule? If the mapping is enumerable and stable, write it. A lookup table with 400 correct entries beats a model with 400 features, and you get determinism, unit tests and zero training cost.
  2. What shape is the input? Named columns in rows means classical ML, and specifically a boosted ensemble first. Raw pixels, audio or free text means deep learning, starting from a pretrained model rather than a fresh one.
  3. How many labelled examples do you have today? Under a thousand: classical models, or a pretrained encoder with a linear head bolted on. One thousand to a hundred thousand tabular rows: boosting, comfortably. Beyond that with rich raw signal, a deep model starts to justify itself.
  4. Is the output a decision or an artefact? A decision — approve, route, score, rank — is discriminative work. Content is generative work, and it comes with a review path you have to design and staff.
  5. What are the latency and money budgets per request? Sub-ten-millisecond responses at high request rates on CPU-only hardware eliminate most of the right-hand columns before quality is even discussed.
  6. Does anyone have to justify the answer to the person it affected? Credit, hiring, insurance, clinical triage, anything appealable — then you need named features and a readable path, and that constraint alone frequently ends the debate in favour of classical ML even where a network would score a point or two higher.

Notice how little of this is about which technique is cleverer. Every question is about something you either possess or do not.

The price of choosing in the fashionable direction

Two versions of the same error account for most of the waste, and both involve reaching upward.

A neural network on 5,000 rows of tabular data. The accuracy loss is the smallest part of the damage. You also acquire a preprocessing pipeline — scaler statistics, encoder vocabularies, imputation defaults — that must be reproduced exactly at serving time or predictions silently degrade, plus a hyperparameter surface with no good defaults where the gap between 0.67 and 0.71 is a learning-rate schedule. Retraining stops being a keystroke and becomes a job with a queue. And when someone asks why an application was declined, the answer is a saliency plot. The ensemble that beat it trains in under a second, needs no preprocessing, and answers the "why" in the language of the business.

A model call for something a lookup table already answers. Ticket routing on subject lines is the canonical example, and it is often genuinely this simple:

Text
import reROUTES = {"refund": "billing", "invoice": "billing", "gst": "billing",          "password": "identity", "otp": "identity", "login": "identity",          "delayed": "logistics", "tracking": "logistics", "courier": "logistics"}WORD = re.compile(r"[a-z]+")def route(subject, default="triage"):    for w in WORD.findall(subject.lower()):        if w in ROUTES:            return ROUTES[w]    return defaultprint(route("Refund not received for order 8812"))   # billingprint(route("OTP never arrives on login"))           # identityprint(route("Your product changed my life"))         # triage

That resolves in microseconds, costs nothing, answers identically every time, and fails into an explicit triage queue. Replace it with a model call and you take on a round trip of hundreds of milliseconds, a per-call charge multiplied by ticket volume, non-determinism on inputs whose answers you already knew exactly, a new failure mode where the returned department does not exist in your system, a vendor dependency, and a regression risk every time the model behind the endpoint changes.

The converse deserves saying, because the boundary is the whole point: as soon as the input is genuinely open-ended — a long complaint with no reliable keywords, mixed languages, three issues in one message — the keyword table falls to triage constantly and the model is worth every millisecond. The mistake is never using the powerful tool; it is paying its maintenance cost for a capability the problem never needed.

Reaching downward is a real error too, just rarer. Hand-crafting spectral features for a speech task, or shape descriptors for a vision task, is months of work a fine-tuned pretrained model will beat. The rule is symmetric: match the tool to the structure of the data, not to the mood of the industry.

The escalation ladder, and when to stop climbing

One sequence works reliably, and its virtue is that every step produces a number the next step has to beat.

Write down the target and one metric before touching a model. Build the dumbest thing that produces a prediction — majority class, one hand-written rule, logistic regression on five features you trust — and record its score. Then fit a gradient-boosted ensemble on every column you have, which takes an afternoon including the feature pipeline. Only when that plateaus below what the product needs do you spend on representation learning, and even then you start from a pretrained model and fine-tune. Skipping the middle step is how teams end up unable to say whether their network beats a decision stump.

Three checks tell you whether the choice was right. Is the gap between the simple and complex option larger than your metric's noise across random seeds and splits? Re-run both with three seeds; if the fancy option's win is smaller than the spread, it has not won anything. Does the expensive option fit the latency and cost budget at peak traffic, not average? And can you say what the model is looking at, in enough detail to answer a complaint from the person the prediction was about?

The test for whether a decision was actually made is whether someone can finish this sentence without hedging: "we are using X because the data is Y and we have Z labelled examples." When nobody can say why a simpler model was rejected, no choice was made — a default was inherited, and defaults in this field are set by whatever was in the news.