Introduction to Generative AI

Key Applications (text, image, music, code)


Four requests land on your desk in one morning. Marketing wants 2,000 product descriptions written from a spreadsheet of specifications. Design wants concept art for a campaign. Audio wants a thirty-second background track for a demo video. Engineering wants 40,000 lines of a legacy accounting system translated into a modern language.

Every one of those is something generative AI is advertised to do. Three of them will go well. One will consume a quarter of engineering time and be quietly abandoned. Guess which, and more importantly, work out why before reading on.

The answer has nothing to do with which domain is "harder" for the model. It has to do with three properties of the task:

  • How expensive is checking? A human can skim a product description in five seconds. Verifying that 40,000 lines of translated accounting logic still computes VAT correctly takes months.
  • How much variation is acceptable? Concept art benefits from ten different options. A payroll calculation has exactly one right answer.
  • Is a good draft worth anything? A first draft of copy saves real time. A first draft of code that compiles but is subtly wrong is worse than no code, because it looks finished.

The legacy migration is the one that fails. Not because code generation is weak — it is one of the strongest applications — but because bulk translation of business-critical logic scores badly on all three axes at once. Hold that framework; it explains the shape of every domain below.

Generative AI succeeds where verification is cheap, variation is welcome, and a draft has value. It fails where any one of those is missing, regardless of how impressive the demo looked.

One sampling trick, four different mediaSample thenext pieceText: fluentdrafts, invented factsImages: concepts,not finished artMusic: texture,weak long structureCode:boilerplate, plausible bugs
Every domain fails in the same shape — confident output where the model had no grounding — but each one hides it differently.

Text

Text is the most mature application, for a structural reason: language is the domain where the training data is largest, the output is fastest for a human to check, and near-misses are still useful.

The mechanism is next-token prediction. The model holds a running context, produces a probability over every possible next fragment, samples one, appends it, and repeats. Everything from a chat reply to a legal summary is that loop under a different prompt.

Where it works well

ApplicationWhy it fitsWatch for
SummarisationSource text is present in the prompt, so the model is compressing rather than recallingQuiet omission of the one clause that mattered
Rewriting and tone adjustmentContent is given; only surface form changesDrift in meaning when the tone shift is large
Structured extraction (text → JSON)Output is machine-checkable against a schemaInvented values for fields that were simply absent
TranslationEnormous parallel corpora; fluency is the main criterionLow-resource languages; idiom; formal register
Drafting from a briefA human edits before anything shipsGeneric phrasing that reads as machine-written
Conversational assistanceThe user is the verification loop, turn by turnConfident answers on questions outside its knowledge

Where it fails

The failure mode has a name — hallucination, meaning fluent invented content — and one cause. Training rewards likely text, not true text. A fabricated case citation is high-probability because it has the exact shape of a real one: plausible party names, a plausible reporter, a plausible year. Nothing in the model's objective separates that from a real citation.

The practical consequence: any text application that depends on facts needs the facts supplied in the prompt, not recalled from the weights. This is retrieval-augmented generation, and it is the single most common architecture in production text systems. Fetch the relevant documents, put them in the context, ask the model to answer using only those, and require it to quote. You have converted a recall problem, which the model is bad at, into a reading-comprehension problem, which it is excellent at.

Images

Most image generators work differently. There is no left-to-right sequence. Instead the model starts from pure random noise and refines the entire canvas over a few dozen steps, each step removing a little of the estimated noise, guided by a text prompt. That is diffusion; many recent open models, such as Stable Diffusion 3 and FLUX.1, use a close variant called flow matching that follows straighter paths from noise to image, but the whole-canvas refinement is the same.

Not every system works this way. Some image generators built into multimodal chat models produce the picture as a sequence of image tokens, one after another, like text; OpenAI describes GPT-4o's image generation as autoregressive, unlike its earlier diffusion-based DALL·E. The practical advice below applies to both.

That whole-canvas refinement matters for how you use it. A text model commits to each word as it goes; an image model revises the whole picture repeatedly, which is why global composition tends to come out well and why fine local detail is where the errors concentrate.

What the modes are

ModeInputTypical use
Text-to-imageA promptConcept art, stock imagery, mood boards
Image-to-imageAn image plus a promptRestyling, variations on an approved direction
InpaintingAn image plus a maskRemoving an object, replacing a background
OutpaintingAn image plus a larger canvasChanging aspect ratio without cropping
Structural conditioningAn image plus a pose, depth map or edge mapKeeping layout fixed while changing appearance
Subject personalisationA handful of photos of one subjectConsistent product or character across a campaign
Restoration and upscalingA degraded imageArchival work, low-resolution assets

Structural conditioning is the one people underuse and professionals rely on. Plain text-to-image gives you no control over composition — you re-roll until something lands. Feeding a rough sketch, a depth map or a pose skeleton alongside the prompt fixes the geometry and lets the model handle only the appearance. That converts an unpredictable slot machine into a repeatable tool.

Where it fails

Three failures recur. Hands and small anatomy — much improved but still the first place to look. Text inside images — recent models render short headlines and signs far better than earlier ones did, but long passages, small print and unusual words still come out garbled, because letterforms demand an exactness that the model can treat as texture. Counting and spatial relations — "three red cubes to the left of two blue spheres" reliably comes back with the wrong count, because the prompt is encoded as a soft bundle of concepts rather than a parsed structure.

Restoration carries a subtler risk. An upscaler does not recover detail that was lost; it invents plausible detail. On a holiday photo that is fine. On CCTV footage or a medical scan it is an evidentiary disaster — the model will happily produce a sharp, confident face that belongs to nobody.

Enhancement is generation. Anything an upscaler adds is a guess, and it will look exactly as convincing as the parts that are real.

Music and audio

Audio is the domain where the gap between "sounds impressive" and "usable in production" is widest, and the reason is structural. A convincing eight-second clip is not evidence of a convincing three-minute piece, because music depends on structure at a timescale far longer than any local pattern.

Modern systems mostly work in two stages. A neural codec compresses raw audio — 44,100 samples per second — into a much slower stream of discrete tokens, perhaps 50 per second. A sequence model then generates those tokens, and the codec decodes them back into waveform. The compression is what makes the problem tractable at all.

ApplicationMaturityPractical limit
Speech synthesis from textHigh — routinely indistinguishable from a recordingEmotional nuance in long-form narration
Voice cloning from a short sampleHighConsent and impersonation risk, not technical quality
Short background and stock musicGoodLoops and beds, not compositions with an arc
Sound effects and foleyGoodPrecise synchronisation to picture
Stem separation and remixingGoodArtefacts on dense mixes
Full-length songsImproving fast — commercial tools now produce complete songs of several minutes with verses and chorusesFine control over arrangement and development; editing one part without changing the rest

The lopsidedness is worth understanding. Speech is essentially solved because the mapping from text to sound is tightly constrained — there are only so many ways to say a sentence. Music is harder because the target is a global shape, and a model predicting the next 20 milliseconds has no built-in representation of "this is the bridge, and it must resolve back to the chorus." Current song generators get much of that shape from lyrics and section markers such as "verse" and "chorus" in the prompt, which is why they manage pop-song structure far better than a piece that develops one idea over ten minutes.

Voice cloning is the application where the binding constraint stopped being technical some time ago. A few seconds of reference audio suffices. The engineering questions that remain are consent, provenance and watermarking — and if you build in this space, they are your questions whether you want them or not.

Code

Code generation uses the same next-token machinery as prose, on text that happens to be a programming language. What changes completely is the acceptance criterion.

Prose degrades gracefully: an awkward sentence in a good paragraph is still readable. Code does not degrade. An off-by-one error in an otherwise perfect function is a bug, and it is a bug that looks exactly like working code. The output is fluent and the requirement is exact, which is the most dangerous combination in the whole field.

Where it earns its keep

TaskValueWhy
Inline completion while typingVery highYou read every line as it appears; rejection costs one keystroke
Boilerplate and scaffoldingVery highHighly patterned, low-stakes, instantly verifiable
Test generationHighTests are checkable by running them; broad coverage is the goal
Explaining unfamiliar codeHighReading comprehension, the model's strength
Small, well-specified functionsHighFits in one screen; a test proves it
Mechanical refactors across a repoMediumDepends entirely on test coverage
Bulk translation of business-critical logicLowVerification cost exceeds rewrite cost

That last row is the morning's fourth request. The model can translate the syntax competently. What it cannot do is guarantee that a decade of accumulated edge cases — the rounding rule for one currency, the fiscal-year boundary someone patched in 2009 — survived. And nobody can verify that quickly, because the specification was never written down; it exists only as the behaviour of the old system. Generated code is cheap. Confidence in generated code is not, and confidence is the thing you were actually buying.

The value of generated code is bounded by your ability to check it. Where the test suite is strong the leverage is enormous; where it is absent the generated code is a liability that compiles.

Much code generation now happens through coding agents: tools that read a repository, edit several files, run the tests and fix what fails, over many steps. They move the "mechanical refactors" row up, because the agent does the tedious run-fix-rerun loop itself. They do not rescue the legacy migration. An agent can only aim at the checks it is given, and the missing piece there is a specification nobody wrote down.

Two specific traps

Invented dependencies. Models produce package names and function signatures that fit the pattern of real ones but do not exist. This has become an attack vector: adversaries register the commonly-hallucinated names on public registries. Check that every import resolves to a package you meant to use.

Insecure defaults. The model reproduces the statistics of its training corpus, and public code is full of string-concatenated SQL, disabled certificate checks and hardcoded secrets. Generated code inherits those habits at roughly the rate they appear in the wild. Static analysis in the pipeline is not optional.

Comparing the four

TextImageAudioCode
Dominant approachAutoregressive transformerIterative denoising (diffusion or flow matching)Codec tokens plus a sequence modelAutoregressive transformer
Generation orderSequentialWhole canvas, refinedSequentialSequential
Cost to verify one outputSecondsInstant — you look at itReal time — you must listenMinutes to weeks
Tolerance for variationMediumHigh — variation is the pointMediumVery low
Main failureFluent falsehoodLocal detail, counting, text-in-imageNo long-range structurePlausible bug
Safe patternRetrieve, then generate, then citeGenerate many, human selectsGenerate short, human assemblesGenerate, then test, then review

Where the domains meet

The commercially interesting systems increasingly span media, and they do it by projecting everything into a shared representation — text, pixels and audio all encoded as vectors in one space, so a model can attend across them.

  • Multimodal language models take images (and often audio or documents) alongside text and answer in text: reading a chart, describing a photograph for accessibility, extracting fields from a scanned invoice. This is no longer a separate kind of model; most leading chat models accept images as standard.
  • Video generation extends image denoising with a time axis, and adds a constraint that images never had — an object must stay the same object from frame to frame.
  • Text-to-3D produces meshes or radiance fields for games and product visualisation, usually by generating many consistent views and reconciling them.
  • Speech-to-speech systems skip the intermediate transcript entirely, which preserves tone and cuts latency enough for natural conversation.
  • Agentic systems chain generation with tool calls — the model writes a query, runs it, reads the result, and decides what to do next. This is where code generation quietly became infrastructure rather than a feature.

What this means when you build something

Score any proposed application on the three axes from the opening before writing a line of integration code.

Verification cost. How long does it take a competent person to decide whether one output is acceptable? If the answer is more than a couple of minutes and there is no automated check, the economics do not work — you will spend more on review than you saved on production. Build the automated check first if one is possible: a JSON schema, a test suite, a linter, a retrieval source to cite against.

Cost of a wrong output that ships. Multiply by the rate at which wrong outputs slip through review, which is never zero and rises as reviewers get comfortable. A generic marketing paragraph costs nothing. A wrong dosage, a wrong tax rate or a wrong legal citation costs everything. High-consequence outputs need a hard gate, not a nudge.

Whether variation helps or hurts. If ten different outputs are ten useful options, put the human at the end and show them all — this is why image tools return grids. If ten different outputs are nine bugs and one correct answer, constrain hard: low temperature, tight schema, validation, retry.

Run the morning's four requests through that grid and the outcome is no longer a guess. Product descriptions: cheap to check, variation welcome, draft valuable — ship it. Concept art: instant to check, variation is the entire point — ship it. A thirty-second music bed: short enough that long-range structure never becomes the problem — ship it. Forty thousand lines of accounting logic: verification cost unbounded, variation fatal, and a draft that looks finished is actively dangerous. That is not a model limitation you can wait out. It is a property of the task.