Course Content
Generative AI System Design Interview
11 sections · 27 lessons
Image captioning: vision-language models, contrastive pretraining, serving and safety
With the purpose, the metrics and the data settled, the rest of the captioning design is where this section earns its place as the bridge of the course. The model introduces cross-attention and vision-language models, the training step introduces contrastive image-text pretraining — the idea Sections 9, 10 and 11 all build on — and serving turns out to have a cost profile unlike any text system so far.
The model
Two architectures, a decade apart, and the second is why this part matters to Sections 9 through 11.
The classical design: vision encoder plus text decoder
The vision encoder turns the image into a sequence of feature vectors. A convolutional network produces a grid of features, each corresponding to a region of the image; a vision transformer splits the image into fixed patches — 16×16 pixels is a common choice — embeds each, and processes them like tokens. Either way, an image becomes something in the region of 200 to 600 vectors.
The text decoder generates the caption autoregressively, exactly as in Autoregressive generation, explained plainly, with one addition: cross-attention. At every layer, while generating each word, the decoder attends over the image's feature vectors and pulls in what it needs.
This is the mechanism that makes captioning work, and it is worth stating precisely: while producing the word "holding", the decoder's attention concentrates on the image regions containing hands and the held object. While producing "beach", it spreads over the background. The image is not compressed into one vector and handed over once; it stays available, and each generated word queries it afresh.
That is also exactly the mechanism Section 9 uses in the other direction, where a diffusion model attends over text embeddings while generating an image. Learn it once here.
The modern design: a vision encoder bolted onto a language model
The shift that changed the field: rather than training a decoder to caption, take a large pretrained language model and give it eyes.
- Take a vision encoder that was pretrained to align images with text (see the training step below).
- Train a small projection — often a couple of layers, sometimes a small transformer — that maps the image's feature vectors into the language model's token embedding space.
- Feed the projected image vectors to the language model as if they were tokens, followed by a text instruction, and let it generate.
Frequently the vision encoder and the language model are both frozen and only the projection is trained, which is remarkably cheap — hours on a small number of accelerators rather than weeks.
What this buys is a change in kind, not degree:
- Open vocabulary and world knowledge. The language model knows what a theremin is, and what a specific architectural style is called, because it read about them.
- Instruction control. The same model produces "describe this for a blind user in one sentence", "list every object", or "what is unusual here?" — one system serving all three products from the framing lesson.
- Reasoning about the image, which is what turns captioning into visual question answering.
What it costs is size. A classical captioner might be small enough to run on a CPU. A vision-language model is a large model with an image occupying hundreds of context tokens, and the serving step below prices it.
Which to choose
| Classical encoder-decoder | Vision-language model | |
|---|---|---|
| Size and cost per image | Small — cents per thousand | Large — see the serving step below |
| Vocabulary | Limited to training captions | Open |
| Controllable by instruction | No | Yes |
| Training cost | Full training run | Often only the projection |
| Best for | High-volume single-purpose captioning at fixed style | Anything needing flexibility, detail, or reasoning |
For a 500-million-image catalogue backfill with one fixed caption style, the classical design is defensible on cost alone. For everything else, the vision-language model wins, and it wins by enough that "we would use a pretrained vision-language model and train a projection" is the expected answer.
Training
Two things to cover: the objective and its known flaw, and the pretraining idea that Sections 9 to 11 all build on.
The objective, and exposure bias
Standard next-token cross-entropy over the caption, conditioned on the image, with teacher forcing — the decoder is fed the correct previous words during training rather than its own predictions, which makes training parallel across positions.
Exposure bias is the gap this creates. At training time the model has only ever seen perfect prefixes. At generation time it sees its own output, mistakes included, and it has never been trained on how to continue from an error. One wrong word early can pull the rest of the caption after it, which is why bad captions often fail as a coherent-sounding whole rather than as one wrong word.
Three responses, honestly rated:
- Scheduled sampling — during training, sometimes feed the model's own prediction instead of the truth, increasing that probability over the run. Helps, and complicates training.
- Sequence-level training with a metric reward — after supervised training, optimise the whole generated caption against a caption metric such as CIDEr, using the model's own greedy output as the baseline to compare against. Historically this produced large gains on captioning benchmarks, and it also drives the generic-caption failure from the metrics lesson, so it must be paired with a semantic metric.
- Scale. With large pretrained vision-language models, exposure bias is much less visible in practice, because the language model is fluent enough to recover. This is the pragmatic modern answer and it is worth saying plainly rather than reciting techniques that are less used than they were.
Contrastive image-text pretraining
This is the most reused idea in the second half of this course. Learn it here and Sections 9, 10, and 11 become much easier.
The setup. Take a large batch of image-text pairs — say 32,000 of them. Encode every image with an image encoder and every text with a text encoder, into the same embedding space.
The objective. For each image, its matching text should be the nearest of all 32,000 texts; for each text, its matching image should be the nearest of all the images. Every other pair in the batch acts as a negative example.
What that produces. An image encoder and a text encoder whose outputs live in one shared space, where "a photograph of a red umbrella" sits near photographs of red umbrellas. Nobody labelled categories; the alignment came entirely from paired data at scale.
Why it matters here and later:
- The image encoder's representation is already language-aligned, so the projection described above has a much easier job.
- It gives you the reference-free image-text similarity metric from the metrics lesson, and the data filter from the data lesson, for free.
- It gives Section 9 its text encoder — the component that turns a prompt into the embeddings a diffusion model conditions on. Text-to-image generation depends on a text representation that is already aligned to images, and this is where that comes from.
- The same in-batch-negatives trick appears in the visual search case study of Machine Learning System Design Interview.
One implementation note that explains the compute bill: quality depends strongly on batch size, because the batch supplies the negatives. Training with tens of thousands of pairs per batch requires distributing the batch across many accelerators and gathering the embeddings — the main engineering difficulty of the method.
Serving
Captioning has an unusual cost profile, and noticing it is worth a point in the interview.
Two paths with different economics
Batch backfill — caption an existing catalogue of 500 million images. Throughput is everything, latency is irrelevant, and the whole job can run on spot capacity overnight for weeks. Batch sizes are as large as memory allows.
On-demand — caption an image at upload. Latency matters (a user is waiting, or an accessibility feature is blocked on it), volumes are lower, and batching is opportunistic.
Design both, and note that they can share a model and share nothing else.
The encoder dominates, which inverts the usual picture
In Sections 2 to 5, decode dominated once outputs were more than a handful of tokens. Here the output is a caption — 15 to 30 tokens — and the input is an image occupying several hundred context tokens after projection.
An illustrative breakdown for a vision-language model on one image:
| Stage | Time | GPU-seconds |
|---|---|---|
| Decode, resize, normalise | 8 ms | CPU |
| Vision encoder | 25 ms | 0.020 |
| Prefill: ~300 image tokens + 60 instruction tokens | 30 ms | 0.030 |
| Decode 25 caption tokens | 28 ms | 0.028 |
| Safety and hallucination checks | 15 ms | 0.008 |
| Total | ~106 ms | 0.086 |
The image side — encoder plus prefill over image tokens — is 58% of the compute. The lever that matters is therefore image resolution, not caption length: halving the number of image tokens (by using a lower input resolution) cuts total cost by roughly a third, and costs detail on small objects. That is the real quality-cost dial in this system.
Caching by image
The same image is captioned more than once more often than you would guess — re-uploads, duplicated catalogue entries, the same stock photograph on many product pages.
- Exact content hash on the decoded pixel data. Cheap, safe, no false positives, and it misses anything re-encoded, resized, or re-compressed — which is most near-duplicates.
- Perceptual hash, which is stable under resize and mild compression. Catches far more, and can collide on genuinely different images — dangerous for accessibility, where a wrong caption is worse than none. Use a conservative threshold and never for images with text in them.
- Embedding-based near-duplicate lookup using the vision encoder's own output, which you computed anyway. Most accurate, and it does not save the encoder pass, only prefill and decode — so it saves about 40% rather than 100%.
An illustrative 15% hit rate on a mixed catalogue is achievable with perceptual hashing.
Cost, computed
At $0.00069 per GPU-second: 0.086 GPU-s × $0.00069 = $0.000059 per image.
- Backfill of 500 million images: about $30,000 as a one-off job. That is a number a team can approve, which is worth knowing before you propose the classical smaller model.
- On-demand at 20 million uploads a day: about $1,180 a day, roughly $430,000 a year, or about $1,000 a day with a 15% cache hit rate.
- The smaller classical captioner, at perhaps 15× less compute, would cost about $2,000 for the backfill and $80 a day on-demand — real money at scale, and the reason the model choice above is a genuine trade rather than a formality.
Safety, monitoring and follow-ups
Two failure modes matter here, and both are characteristic of cross-modal generation rather than incidental.
Object hallucination
The signature failure: the caption confidently names something that is not in the image. A beach scene gets a surfboard. A kitchen gets a knife. A desk gets a laptop.
The cause is worth understanding, because it explains the fix. The decoder is a language model, and language models have strong priors about what words follow other words. Having generated "a man standing on a beach holding a", the language prior for "surfboard" is high, and if the image evidence at that step is weak, the prior wins. Hallucination here is the language half of the model overruling the vision half.
Four mitigations:
- Grounding verification. Run an open-vocabulary object detector over the image, extract the nouns from the caption, and check each has a detection above threshold. Nouns that fail are removed or the caption is regenerated. Costs about 15 ms and catches the common cases.
- Lower decoding temperature. Captioning is a task with a right answer (see Decoding strategies), so greedy or low-temperature decoding is correct and reduces the chance of a low-evidence token being sampled.
- Prompt for hedging. Instructing the model to describe only what is clearly visible measurably reduces confident invention, at some cost in fluency.
- Abstain. For accessibility especially, "an image that could not be described reliably" is better than a wrong description, because the user cannot check it. Design abstention as a product behaviour, exactly as in Generation grounded in context.
Describing people
This is where a captioning system does the most harm, and it needs an explicit rule rather than a general intention.
The rule: describe what is visible, never infer an attribute. A caption should say "a person in a red jacket", not "a young woman in a red jacket". Gender, age, race, nationality, religion, disability, and occupation are not visible properties; they are inferences, and a model trained on web captions makes them the way its training data did.
Two supporting reasons, and both should be said out loud. It is an accuracy requirement — the inference is frequently wrong. And it is a dignity requirement — image-labelling systems have produced grotesquely racist outputs in well-documented public incidents, and the mechanism was exactly this: a model asked to assign categories to people, trained on data with a skewed distribution, with no constraint stopping it.
Implement it as a constrained-decoding and post-filter rule, not as a prompt instruction alone. Maintain a list of attribute terms that may not be applied to people, strip them, and regenerate. A prompt is a preference; a filter is a guarantee.
Evaluating quality across groups
Aggregate caption quality hides per-group failure, and the failure is rarely uniform.
Build a stratified evaluation set: images of people across skin tones, ages, apparent disabilities, and geographies, and scenes from a wide spread of countries, all with human references. Then report accuracy, hallucination rate, and specificity per stratum, and treat a gap between strata as a launch blocker rather than a note.
The practical difficulty is that building such a set requires collecting images of people with consent and demographic metadata, which is itself sensitive data with retention obligations. Budget for it; there is no shortcut, and shipping without it means finding out from users.
The extensions interviewers ask about
- Dense captioning — many region-level captions rather than one global sentence. Better for complex images and far more expensive: it is a detection problem plus one caption per region.
- Visual question answering — the same architecture with a question in the prompt. The modern design described above gets this nearly for free, which is one of its strongest arguments.
- Video captioning — sample frames, caption them, and summarise across time, or use a video encoder. Section 11 (Text-to-Video Generation) covers why temporal modelling is harder than it looks.
- Interactive captioning — the user asks a follow-up about the image. Combines this section's design with the chatbot's context management.