What is generative AI? Picking an answer versus composing one

JR

Jai Rao

August 22, 202617 min read

A beginner guide to generative AI: how composing output differs from picking a label, what text, code, image and audio models do well, and why hallucination is structural.


For most of the last decade, the AI systems that actually ran in production did one job: they sorted things. Is this email spam or not spam. Is this transaction fraud or fine. Does this photo contain a pedestrian. Whatever came out was one item pulled off a list that an engineer had written down in advance, and the entire system was built around getting that pick right more often than the last version did.

Generative models do something else, and the difference is not "the same thing, but better". They produce artefacts that did not exist before the request: a paragraph, a working function, a photograph of a room nobody has ever photographed, a voice reading a sentence nobody has ever said. Nothing is being selected off a list. Something is being assembled. That change in the kind of output is what the word "generative" is pointing at, and nearly everything that surprises people about these systems, including the specific ways they fail, follows from it.

A classifier picks; a generator composes

Start with the classifier, because it is the simpler machine and the contrast is where the whole idea lives.

A sentiment classifier has three possible answers: positive, negative, neutral. Somebody chose those three. The model reads a review and produces three numbers that add up to one, and you take the largest. A big image classifier might have a thousand labels instead of three. That is a much longer list, but it is still a list, and a human wrote it down before the model ever ran. Which means you can always check the answer: there is a correct item, and the model either picked it or it didn't.

Now try to build the same machine for writing English. What is the list? Suppose the model works in pieces of text drawn from a fixed vocabulary of roughly fifty thousand pieces, and you want two hundred pieces of output. The number of possible outputs is fifty thousand raised to the power of two hundred: a one followed by something like nine hundred and forty zeros, against perhaps ten to the eightieth atoms in the observable universe. You cannot score every candidate and return the best one, because you cannot write the candidates down.

So the model does the only tractable thing available. It converts one impossible choice into two hundred easy ones. At each step it looks at everything it has so far, meaning your request plus whatever it has already produced, and assigns a probability to every piece in that fixed vocabulary. It picks one. It sticks that piece onto the end of the input. Then it asks the identical question again, with an input that is now one piece longer.

That gives you the sentence worth remembering: a generator is a classifier in a loop, one whose own output becomes part of the next question. Each individual step really is a pick from a fixed set. The finished sequence is on no list at all. The novelty comes from composition, not from some creative faculty tucked away inside the network. It also explains why generation is slow and arrives in a trickle: the two hundredth piece cannot be computed until the one hundred and ninety-ninth exists, so the work is inherently sequential in a way that classification is not.

PropertyClassifierGenerator
What comes outOne label from a fixed setA sequence assembled piece by piece
Size of the answer spaceThree to a few thousand, written down by a personAstronomical, never enumerated
Cost of one answerOne pass through the networkOne pass per piece produced
Checking the answerCompare against the labelled truthNeeds judgement; several answers are often acceptable
What "wrong" looks likeThe wrong item off the listFluent, well-formed, and false
Measuring the modelAccuracy on a held-out setGraded examples plus human review

Learning to write by finishing other people's sentences

Nobody hand-labelled the ability to compose sentences into a model. The training procedure is almost insultingly simple to describe. Take an enormous quantity of text. Cut a document at an arbitrary point, hide what came next, and ask the network to guess it. Compare the guess to the piece that actually followed, then nudge the network's internal numbers so that the real piece would have been rated a little more likely. Repeat an unthinkable number of times.

The reason this scaled to where it did is that the data supervises itself. There is no annotation budget and no labelling team: any text that exists is already a finished training example, because the answer to "what comes next" is sitting right there in the document. That cheapness is the whole story of how these models got as large as they are.

Why does such a narrow objective produce something broader than autocomplete? Because being less wrong about the next piece forces the network to pick up any regularity that helps, and an enormous amount of the world turns out to help:

  • "The capital of France is ___" — the fact has to be stored in the weights somewhere, or the guess is no better than chance.
  • "She tipped the jug and the water ran into the ___" — needs to know that liquids end up in containers, and which containers people actually own.
  • "def total(items): return sum(" — needs a language's conventions, its indentation rules, and the notion that a function called total probably adds things up.
  • "The witness contradicted herself twice, so the jury ___" — needs the shape of an argument, not just the shape of a phrase.

Nothing in the training objective mentions geography, physics, Python, or law. Those appear because they are the cheapest available route to being less wrong about the next word. Prediction turns out to be demanding enough that doing it well requires absorbing the structure of whatever is being talked about. The machinery that carries this out — how text gets chopped into pieces, how each position in a sequence looks back at earlier ones — is a substantial subject in its own right, and you genuinely do not need it to use these systems competently.

Why predicting the next word starts to resemble knowing things

Here is the intuition that makes the leap from "statistics over text" to "behaves like it knows things" feel less like magic. A model has a fixed number of internal parameters. It is a large number, but it is small next to the volume of text the model was trained on. Storing the training data is not an option. What the network can store is what the training data has in common. General rules take far less room than a billion special cases, so under that pressure, general rules are what survive: grammar first, because it is the most repeated regularity there is, then facts, because a fact stated ten thousand ways is cheaper to store once, then the shape of an explanation and the conventions of a code file.

Be precise about what kind of knowledge this is, though, because the imprecision is where people get hurt. The model learned a distribution over text about the world. It never saw the world. Where the text is dense, consistent, and repeated in many different phrasings — capital cities, common idioms, the standard library of a popular language, how a business email is structured — that distribution is sharp, and the output is genuinely dependable. Where the text is thin, contradictory, or simply absent — your company's approval process, one obscure paper, what happened last Tuesday — the distribution is flat.

And this is the part that catches nearly everyone: a flat distribution does not produce flat-sounding prose. The same machinery runs either way. The output arrives in the same confident register, with the same clean grammar, at the same speed. On the surface, there is no difference between "ten thousand consistent sources agreed on this" and "this is the shape an answer to that question would have". You have to know which regime you are in from the outside, because the text will not tell you.

Text in, text out: the entire interface

For all the machinery underneath, the interface is unglamorous, and seeing it once dissolves a lot of mystique. You send a string. You get a string back. Here is the whole thing in Python, using the official SDK.

Text
import anthropicclient = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environmentresponse = client.messages.create(    model="claude-opus-5",    max_tokens=1000,    messages=[        {"role": "user",         "content": "Explain a hash map to a ten-year-old, in three sentences."}    ],)for block in response.content:    if block.type == "text":        print(block.text)

Three things in that snippet are worth noticing. The request carries no state, so the model has no recollection of any previous call — a chat interface is just your program resending the whole conversation every single time. There is a hard ceiling on how much text can come back, set by you. And running it twice can produce two different answers, because each piece is sampled from a distribution rather than read off a table; if you need the same text again, save it rather than regenerating it. Everything else people build around generation — conversation history, document loading, tool calls, checking the output before it ships — is ordinary software sitting on top of that one call.

The four modalities, and what each is honestly good at

Text is the strongest and the most general. It is at its best when the material is already in the request and the job is to change its form: summarise this, rewrite it for a different reader, translate it, pull these five fields out of it, sort these tickets into categories. That last one is a pleasing loop — generators are now routinely used to do the classifier's old job, because writing the categories in a sentence is faster than assembling a labelled dataset. Text generation is at its weakest when the model is being asked to serve as an authority on facts it has to dredge up from memory.

Code is arguably the best-fitting modality of the four, for reasons that have nothing to do with code being easy. Code exists in enormous public volume, it is heavily patterned, and — this is the decisive one — it is verifiable. You can run it. The feedback loop closes without anyone making a judgement call. Expect real value from boilerplate, glue between two libraries, test scaffolding, translating a routine from one language to another, and explaining unfamiliar code. Expect trouble when a large design has to stay coherent across many files, and expect the model to produce, with total composure, a function call in exactly the right style for a version of a library in which that function never existed.

Images take a text description and return a picture. The mechanism differs from text generation: rather than extending a sequence left to right, these models start from noise and repeatedly remove a bit of it, steered at each step toward the description. What they are genuinely good at is mood, composition, style, texture, concept exploration, and above all the speed of iterating on a look. What they are unreliable at is exact text inside the image, precise counts of objects, rendering the same character consistently across two pictures, and anything where spatial relationships have to be exactly right — a wiring diagram, a floor plan, a chart with correct values.

Audio is really two different jobs. Speech synthesis is effectively solved for most purposes; what remains hard is prosody over a long passage and taking direction on emotion. Speech recognition is strong on clear recordings in well-represented languages, and degrades on overlapping speakers, heavy background noise, accents thin in the training data, and domain jargon. Music generation produces convincing texture and convincing pastiche, but holding a structure together over several minutes is far harder than making thirty seconds sound right.

One thing cuts across all four: each is at its best when a person reads the output and a bad one costs a wasted minute, and at its worst when the output flows straight into a consequential decision that nobody looks at.

Hallucination is the mechanism working, not the mechanism failing

The word is unfortunate, because it suggests a malfunction — some bug a future release will patch out. That is not what is happening, and understanding why is the highest-value idea here.

The model has exactly one behaviour: extend the input with a plausible continuation. There is no separate mode called "answer truthfully" and no separate mode called "invent something". Both are the same operation running through the same weights. When the learned distribution strongly supports the true continuation, plausible and true land in the same place and you get a correct answer. When it doesn't, the plausible continuation is still available, and it comes out with the same fluency, the same grammar, and the same self-assurance as a correct one.

Academic citations are the cleanest illustration. A model has seen tens of thousands of references and has learned their shape exhaustively: two or three surnames, a year, a title built from the right sort of noun phrases, a journal name, a volume, page numbers. Producing a string with precisely that shape is easy, because it is one of the best-attested patterns in written English. Producing one that names a paper which actually exists is a completely different and much harder demand — it requires that specific paper's details to be both stored and retrievable on cue. So you get plausible authors, a title that sounds exactly like a real title, and an identifier that resolves to nothing. Nothing broke. You asked for a continuation and received an excellent one.

This is also why "are you sure?" is close to worthless as a check, and why a model that apologises and revises has told you nothing about the truth of either answer. Hedging is itself a learned continuation. It appears where the surrounding text pattern calls for hedging, not where the model is actually short of evidence. In a person, hesitancy carries real information about their confidence, which is precisely why the mismatch here catches capable people off guard.

What genuinely lowers the rate is structural, not verbal: put the material in the request instead of asking the model to recall it; keep the task closer to transformation than to recollection; ask for output whose correctness you can check mechanically; and place the verification step outside the model, since a system with one behaviour cannot audit itself.

Arithmetic: reciting versus computing

Ask a model what seven times eight is and you get fifty-six. Ask it for four thousand nine hundred and thirteen times two thousand two hundred and seven and you will very likely get a number that is wrong. The gap between those two cases is not a gap in effort or model size. It is the difference between reciting and computing.

Digits come out the same way words do: one piece at a time, from pattern. "Seven times eight is fifty-six" appears constantly in text, so it is recited, and recitation is reliable. A product of two specific four-digit numbers almost certainly never appeared anywhere in the training text. Getting it right requires executing an algorithm — partial products, carries, no drift across several steps — and then emitting the answer from the most significant digit first, which is close to the opposite of the order in which long multiplication actually settles its digits.

So what you get instead is a number with the right digit count and a believable leading digit that is wrong somewhere in the middle. That is the worst possible failure shape: wrong, but shaped like right. It generalises well past multiplication — counting items in a long list, date arithmetic, sorting exactly, applying a bracketed fee schedule, checking whether a column of figures adds up. Anything with exactly one correct answer reachable only by following a procedure.

The fix is not a bigger model. It is a calculator. This is what tool use is for, and every serious production system does it: the model decides what needs computing and writes the expression, a real interpreter evaluates it, and the result comes back into the conversation as text. Use the generator for the judgement of what to do, and deterministic code for the doing.

Three blind spots that come with the architecture

The first blind spot is the model's relationship to its own ignorance. There is no internal flag that fires when the requested information is not in there. The model produces a distribution over next pieces, and a flat distribution and a sharp one both get turned into fluent sentences by exactly the same procedure. A model can be wrong and definite, and nothing in the output separates that from being right and definite. A classifier, by contrast, at least hands you a probability you can threshold.

Frozen at training time

The weights stop changing when training ends, and nothing that happened afterwards is in them. The model cannot learn from your conversation either: your corrections shape the rest of that conversation because they sit there in the input, and they evaporate when it closes. When a chat product appears to know today's news, something outside the model fetched text and placed it in the input; the model is reading, not remembering. Even asking a model where its own knowledge stops gives an unreliable answer, because it infers that boundary from patterns in text rather than reading a configuration value.

Sources it never saw

Attribution from memory is just generation with a bibliography's formatting. A model can reliably tell you where a claim came from only when the source is in front of it in the request. Otherwise it produces a reference-shaped string that satisfies every surface expectation and may point at nothing. Treat every unrequested citation, link, quote, page number, and version number as a search query rather than a fact. Checking costs seconds, and the consequences of not checking have been demonstrated publicly and repeatedly.

Where to start, and how to tell when it's the wrong tool

The useful question is never "is this thing smart". It is whether the shape of your task matches the shape of the machine. Tasks that fit well share a family resemblance:

  • The material is already in the request, and the job is to change its form rather than retrieve facts.
  • You can check the output faster than you could have produced it yourself.
  • Several different answers would be acceptable, so there is no single string to hit.
  • A bad output costs a wasted minute, and a person sees it before it matters.

Tasks that fit badly are just as recognisable, and recognising them early saves a great deal of wasted effort:

  • Exactly one answer is correct and reaching it requires a procedure — arithmetic, exact counting, sorting, precise comparison.
  • Correctness depends on something recent, private, or specific that you have not supplied in the request.
  • The output feeds an irreversible action with nobody reading it in between.
  • You need to know when the system is unsure, and you need that signal to be trustworthy.

To build the instinct rather than read about it, run two experiments. Take a task where you already know the answer cold: summarise a document you wrote yourself, or ask for a function you have written before. Run it five times and read all five outputs. You will see how much they vary, and that all five sound equally assured. Then take something just past the edge of your own knowledge and check every specific — every name, number, function signature, and date. Some will be correct. Some will merely be shaped like correct.

The distance between how those two exercises feel, which is identical, and how they turn out, which is not, is the thing worth carrying away from this article. People who get real value out of generative models are rarely the ones with the cleverest phrasing tricks. They are the ones who can tell within a few seconds which of those two situations they are in, and who put a check in place for the second one before it costs them something.