Introduction to Generative AI

What is Generative AI?


Suppose a friend asks you for a program that writes a fresh bedtime story every night. No repeats. Try to build it without any machine learning and watch what happens.

Attempt one: templates. You write "Once upon a time, a {ADJECTIVE} {ANIMAL} lived in a {PLACE}." and fill the slots from word lists. It works for about four nights. By night five the child notices that every story has the same skeleton, and a brave rabbit in a castle is not meaningfully different from a clever fox in a forest. The sentences are new; the stories are not.

Attempt two: a big library. You scrape ten thousand real stories into a database and serve one at random. Now the stories are good — but they are not new. You are a librarian, not an author. And on night ten thousand and one you are out of stock.

Attempt three: cut and paste. You chop the library into sentences and recombine them. The output is new in the sense that nobody wrote that exact sequence before, and it is also incoherent: the dragon dies in paragraph two and orders breakfast in paragraph four. Novelty without structure is noise.

What you actually want sits between attempts two and three. You want a machine that has read the ten thousand stories, extracted what stories are like — how they open, how tension builds, which words follow which — and can then produce a brand-new one that obeys those regularities without copying any particular source. That capability is what the word generative names.

Predict a distribution, then roll the diceReadeverythingso farScore everypossiblenext tokenSample onefrom thatdistributionAppend itand repeatTake the highest-scoring token every time and the same prompt gives the same story forever.
Novelty comes from the sampling step, not from the network — the model itself is deterministic.

What "generative" actually means

A generative model is a system that has learned the distribution of some kind of data and can draw new samples from it.

"Distribution" is the piece of jargon to unpack. A distribution is just a statement about which things are common and which are rare. If you looked at every photograph of a face ever taken, faces with two eyes would be overwhelmingly common, faces with one eye vanishingly rare, and images of random static essentially absent. Written as maths, a distribution assigns a probability p(x)p(x) to every possible item xx — every possible photograph, every possible sentence, every possible melody.

Almost all of that probability mass sits in an astonishingly thin slice of the space. Consider a small 256×256 colour image. The number of possible pixel arrangements is 256256×256×3256^{256 \times 256 \times 3}, a number with roughly 470,000 digits. The overwhelming majority of them look like television snow. The images a human would call "a photograph" occupy a sliver so small it defies intuition. A generative model's entire job is to learn where that sliver is, and then to reach into it.

Learning to generate means learning what makes something plausible — and plausibility is a statement about probability, not about truth.

Once a model holds an approximation of p(x)p(x), generation is sampling: rolling dice weighted by that learned probability. Roll once, get a story. Roll again, get a different story. Both are plausible because both came from the high-probability region; both are new because the dice did not have to land where they landed last time.

Unconditional and conditional generation

There are two flavours, and the difference matters in practice.

Unconditional generation samples from p(x)p(x) with no steering: "give me any plausible face." Conditional generation samples from p(x∣c)p(x \mid c) — the distribution of xx given some condition cc: "give me a plausible face, given that the person is elderly and smiling", or "give me a plausible next word, given the 400 words that came before."

Nearly every system people actually use is conditional. A chat assistant is sampling from "plausible reply, given this conversation". An image tool is sampling from "plausible image, given this text prompt". The condition is the steering wheel; without it you get output that is well-formed and completely unrelated to what you wanted.

Recognising versus producing

For most of machine learning's history, the goal was to recognise: sort email into spam and not-spam, read a postcode, spot a tumour. These systems answer a question about an input they are handed. A generative system produces an output nobody handed it. That difference cascades into almost everything else about how the two are built and judged.

AspectRecognition (discriminative)Generation
Question asked"What is this?""What would a plausible one look like?"
Input → outputRich input → small labelSmall input (prompt, noise) → rich output
LearnsWhere the boundary between classes liesWhat the data itself looks like
Correct answerExactly oneEnormously many, all equally valid
ScoringAccuracy against a labelled answer keyHuman judgement, or statistical proxies
Typical failureConfident wrong labelFluent, well-formed nonsense
Data neededExamples with labelsExamples alone; labels optional

That last row is quietly the most consequential. Labelling is expensive — a radiologist has to mark each scan, an annotator has to tag each sentence. Generative training mostly does not need that. Hide part of the data, ask the model to predict it, compare against what was actually there. The data labels itself. That is why generative models could be trained on trillions of words of ordinary internet text while labelled datasets stalled in the millions.

The scoring row is where newcomers get burned. With a classifier you have an answer key: 94% accurate, done. Ask "is this generated story correct?" and the question dissolves. There are billions of good stories. Evaluation of generative systems is genuinely hard, still partly unsolved, and never as clean as a single accuracy number.

The mechanism: predict, then roll the dice

Strip away architecture and almost every generative model does the same two things.

Step one, during training: predict the missing piece. Show the model a fragment of real data with something removed, have it guess what was removed, and adjust its internal numbers when it guesses badly. Repeat billions of times.

  • Text: remove the next word. "The capital of France is ___" → the model must put high probability on "Paris".
  • Images: add random noise to a photo and ask the model which noise was added, so it learns to strip noise away.
  • Audio: predict the next slice of waveform or spectrogram from the slices before it.

Nobody labelled anything. The data supplied its own supervision. This is why the approach is called self-supervised.

Step two, at generation time: sample from the prediction. The model does not output a word. It outputs a probability for every word it knows. Given "The cat sat on the", the numbers might look like this:

Text
mat      0.31floor    0.18couch    0.12windowsill 0.07table    0.05...      (50,000 more, most near zero)

Always taking the top item — "mat" — gives you a model that is deterministic and, in longer passages, dull and repetitive. It also gets stuck in loops. Sampling in proportion to the probabilities gives you variety, which is why asking the same question twice gives different phrasings.

The dial that controls this is temperature. Before sampling, each score is divided by a temperature TT and then renormalised. Work through what that does with the numbers above:

TemperatureEffect on the oddsOutput characterFits
T→0T \to 0The top option takes essentially all the probabilityDeterministic, repetitiveCode, extraction, factual lookup
T=0.7T = 0.7Mild flattening; "mat" still likely, "couch" plausibleCoherent with varietyGeneral assistant use
T=1.0T = 1.0The model's raw learned odds, untouchedNatural, occasionally looseCreative drafting
T>1.4T > 1.4Rare options get inflated towards the common onesSurprising, then incoherentRarely useful

This is the first thing to reach for when output feels wrong in character rather than wrong in content. Rambling, meandering answers usually mean the temperature is too high. Robotic, stuck-in-a-groove answers usually mean it is too low.

One caveat when you use a hosted model. Some current reasoning models, which work through a problem internally before answering, accept only the default temperature or reject the setting altogether. The idea above still describes what happens inside; you just may not get the dial. Check your provider's documentation before building on it.

A generative model does not choose an answer. It shapes a probability landscape, and the sampler decides where in that landscape to land.

The same trick, four different media

The reason one idea covers text, pictures, sound and software is that all four can be cut into a sequence or a grid of small units, and predicting units is a solved problem. What differs is the unit and how much of the output must be coherent at once.

DomainUnit being modelledDominant approachHard part
TextTokens — words or word fragmentsAutoregressive transformer, one token at a timeStaying factual and consistent over long passages
ImagesPixels, or compressed patches of themMostly iterative denoising from random static (diffusion and its close cousin, flow matching); some systems predict image tokens one at a time insteadFine structure — hands, small text, counting objects
Audio and musicWaveform slices or spectrogram framesDenoising, or token prediction over audio codesLong-range structure; a song needs a shape
CodeTokens, same as textAutoregressive transformerIt has to actually run, so "nearly right" is wrong

Code deserves a note. Superficially it is just text, and the same machinery generates it. But the acceptance criterion is unforgiving in a way prose never is. A story with one clumsy sentence is still a story. A function with one wrong variable name is a crash. That mismatch — fluent output, brittle requirements — is why generated code needs review rather than trust.

What generative AI is not

Four confusions cause most of the disappointment people feel with these systems.

It is not a search engine

A search engine finds a document that exists and returns it. A generative model composes an answer from statistical regularities. When it produces a citation, it is producing a plausible-looking citation — an author name that fits, a journal that publishes such work, a year that makes sense. Plausible is not the same as real. This is why hallucination — confident, fluent, invented content — is not a bug that will be patched out. It is the same mechanism that produces good output, applied where the model's knowledge is thin.

Many chat products now run a real web search first and paste the results into the prompt. That helps a great deal, because the model is then summarising documents that exist. But the search is a separate tool bolted on; the model is still composing text, and it can still misquote what it was given.

It is not a database of its training data

A model with a few billion parameters cannot store the terabytes it trained on. It compresses patterns, and compression discards specifics. This has an upside — it can generalise to prompts nobody wrote — and a downside: recall of any particular fact is approximate and gets worse for facts that appeared rarely.

It is not reasoning from principles

When a model produces a correct arithmetic result, it has not necessarily performed arithmetic. It has produced the continuation that fits the pattern of correct arithmetic in its training data. Often that lands correctly. On unusual inputs it lands confidently and wrongly. The same applies to legal reasoning, medical inference, and any domain where correctness comes from rules rather than from resemblance.

Newer reasoning models narrow this gap. They are trained to write out a long chain of intermediate steps before the final answer, and on maths, code and multi-step problems they are much more accurate than models that answer straight away. But the steps are still generated text. They make it more likely the answer follows the rules; they do not guarantee it. Where the answer must be exact, a calculator, a test or a human still has the last word.

It has no notion of truth

The training objective rewards likely, not true. A statement that is wrong but commonly written scores well; a statement that is right but rarely written scores badly. Nothing in the loss function distinguishes the two. Any system that must be right therefore needs an external check — retrieval against real documents, a test suite, a calculator, a human.

These systems are optimised for plausibility. Every safeguard you will ever need to build exists because plausibility and correctness are not the same thing.

What this means when you build something

Three habits follow directly from the mechanism, and they separate people who ship working generative features from people who ship demos.

Design for a distribution of outputs, not one output. Your model will return something different every time. If your product breaks when the wording changes, the product is wrong, not the model. Either constrain the output shape — use the provider's structured-output mode, which makes the model follow a JSON schema you supply, then still validate the values and retry on failure — or build an interface where variation is a feature, such as showing four options and letting the user pick.

Put the check where the cost is. Match verification to the price of being wrong. Draft marketing copy that a human reads before publishing needs no machinery. A generated database query that will run against production needs a parser, a read-only role, and an approval step. The model's confidence tells you nothing about which case you are in — it is equally fluent either way.

Feed the condition, do not hope for it. The model samples from p(x∣c)p(x \mid c), and cc is exactly what you supply — nothing more. If the answer depends on your company's refund policy, that policy has to be in the prompt. Most "the model got it wrong" reports are really "the model was never given the information", and the fix is retrieval rather than a better model.

Hold onto the bedtime-story problem. Templates gave structure without novelty. The library gave quality without novelty. Cut-and-paste gave novelty without structure. What makes a generative model useful is that it holds both at once — and every limitation covered here is the price of that trade.

Check your understanding

0 of 3 answered

1.An assistant cites a research paper with a real-sounding author, journal and year, but the paper does not exist. What is the best explanation?

2.You send the same question twice and get two differently worded answers. Why?

3.A support bot keeps giving wrong answers about your company's refund policy. What should you try first?