Generative AI System Design Interview

Course Content

Generative AI System Design Interview

11 sections · 27 lessons

Image captioning: framing, caption metrics and paired data


Design a system that takes an image and produces a sentence describing it.

This is the bridge section. Sections 2 to 5 were text in, text out. Sections 7 to 11 generate images and video. This one is image in, text out — the first cross-modal system, and the place where the ideas that Sections 9, 10, and 11 depend on are introduced.

Two products behind one modelAlt text for accessibility• One user, who cannot verify the caption• A wrong caption is worse than none• Abstention must be an allowed outputCaptions for search• Indexed in bulk, read by a ranker• A wrong caption costs one bad result• Recall matters more than precision
The accessibility case sets the safety bar, because the reader cannot check the claim.

Clarifying questions

The first one changes the entire product, so ask it first.

What is the caption for? Three answers, three different systems:

  • Accessibility — a screen reader reads it aloud to a blind or low-vision user. The caption must convey what is functionally relevant, in the order that matters, without padding.
  • Search indexing — the caption is never read by a person; it is tokenised into an index. Density and coverage matter; fluency does not.
  • Content understanding — feeding moderation, recommendation, or organisation. Structured attributes may serve better than a sentence.

The other clarifying questions:

  • How detailed? One sentence, or a paragraph? A one-sentence caption of a busy street scene discards almost everything.
  • Which languages? Captioning in one language and translating is cheaper and loses culturally-specific detail; captioning natively per language costs more.
  • What is the image distribution? Product photography, user snapshots, medical imaging, and screenshots are entirely different problems.
  • Batch or on demand? A 500-million-image catalogue backfill and a live upload path have different budgets. Assume both, as the serving lesson does.

The framing

Cross-modal sequence generation: a fixed image in, a variable-length token sequence out.

Structurally this is the translation problem from Section 3 with one substitution — the source is an image rather than a sentence. The encoder-decoder argument from the Google Translate framing lesson applies unchanged: the source is complete and fixed before any output is written, so it should be encoded fully and attended over.

Why the accessibility case deserves naming

If the answer is accessibility, "good" changes meaning in three ways worth stating explicitly.

Relevance beats completeness. "A graph showing revenue rising from 2023 to 2026" is a good caption. "A rectangular image containing a white background, a blue line, axis labels, and a legend" enumerates and communicates less.

Length has a cost. The caption is read aloud in sequence. A user cannot skim it. Twenty unnecessary words is twenty unnecessary seconds across a page of images.

The only competent judge is a user of the assistive technology. Sighted raters systematically prefer captions that read well over captions that are useful, because they can see the image and are grading prose. Recruit blind and low-vision evaluators for this product. It is a small, cheap change to your evaluation plan and it is the one that decides whether the feature is any good.

Metrics

Captioning has more automatic metrics than any other task in this course, which is a symptom rather than a strength: each was invented because the previous one was inadequate, and the newest is still inadequate.

Each metric exists because the last one failedBLEU —n-gram overlapMETEOR —adds synonymsCIDEr —weights rare termsSPICE —scene graphsCLIPScore —no referencetopbottomFive metrics is a symptom, not a strength.
Every rung was added to patch a blind spot in the rung below, and the top one is still inadequate.

The four reference-based metrics

Each compares a generated caption against several human-written reference captions for the same image.

MetricWhat it measuresWhat it rewardsIts blind spot
BLEUWord n-gram overlap with referencesMatching common phrasingsSynonyms; short captions score oddly; barely sees meaning
METEORUnigram matching with stems and synonyms, plus a word-order penaltyParaphrase and correct orderingStill surface-level; needs language resources
CIDErConsensus n-gram similarity, weighting each n-gram by how rare it is across the corpusDistinctive content words — the ones that describe this image and not every imageNeeds many references; still no notion of truth
SPICEParses caption and references into scene graphs of objects, attributes, and relations, and compares thoseSemantic content — did you name the right objects with the right relationsDepends on a parser that makes mistakes; ignores fluency entirely

CIDEr is the one to understand. Its rarity weighting is what stops "a photo of" from earning credit — an n-gram that appears in most captions carries almost no weight, so the score is driven by the words that make this caption specific. That is the closest any of the four gets to measuring whether you described this image.

SPICE is the complement: it ignores how the caption reads and asks only whether the propositions are right. Report CIDEr and SPICE together and you have coverage of distinctiveness and of semantic correctness. Neither knows whether the objects are actually in the picture.

Reference-free image-text similarity

A different instrument: score the caption against the image directly, using a model trained to place matching images and texts near each other in a shared embedding space (see the training lesson).

The advantage is large — no references needed, so it runs on any image including live traffic. The limitations are equally real: it tends to reward captions that mention salient objects and is weak on relations, counting, and negation, and because it is a model it can be gamed by captions that look right to it in ways they do not to a person. Use it as a continuous production monitor, not as the release gate.

Human evaluation, on two axes

Accuracy — is everything the caption says actually in the image? This catches the characteristic failure from Safety, monitoring and follow-ups, where a caption names an object that is not there.

Usefulness — does the caption serve the reader's purpose? For accessibility, this is the only axis that matters and it must be judged by screen-reader users. For search indexing, "usefulness" is measured downstream instead: does adding the caption to the index improve retrieval of images by text query? That is an A/B test, and it is a better metric than any caption-level score.

Data

Captioning needs paired images and text. There are two sources, they differ by four orders of magnitude in size and by a similar margin in quality, and a real system uses both.

Curated pairs against web alt-textCurated datasets• Hundreds of thousands of images• Five human captions per image• Clean, literal, and narrow in subjectWeb alt-text• Billions of image and text pairs• Often a filename or a brand slogan• Broad coverage, heavy filtering needed
Pretrain on the noisy billions, then fine-tune on the clean thousands — neither alone works.

Curated datasets

Human annotators write captions for a fixed set of images, typically five independent references per image. Sizes are in the hundreds of thousands to low millions of images.

Quality is high and consistent. The limitations to name: the images come from a particular collection process and therefore a particular slice of the world; annotator instructions shape the captions in ways that show up in every model trained on them ("describe the main subject in one sentence" produces models that describe one subject in one sentence); and the vocabulary is narrower than reality.

Web alt-text at scale

Billions of image-text pairs harvested from web pages. Enormous, free, and noisy in specific ways worth listing, because the filtering pipeline follows from them:

  • Empty or placeholder — image, img_4021.jpg, photo.
  • Search-engine keyword stuffing rather than description.
  • Filenames and camera metadata.
  • Text about the page, not the image — a caption describing the article the image sits in.
  • Correct but useless — Figure 3.

The standard filter is an image-text similarity model (see the training lesson): score every pair, discard everything below a threshold. It typically removes the large majority of harvested pairs and that is the pipeline working. Add length bounds, language identification, deduplication by image hash, and removal of pairs whose text is boilerplate across many images from the same site.

The practical recipe: pretrain on the large noisy set for coverage and vocabulary, then fine-tune on the small curated set for caption style and consistency. Scale first, quality second.

The bias captions inherit

Captions are written by people about images collected by people, and both steps carry assumptions into the model. Three that reliably appear:

  • Geographic skew. Web image-text data over-represents wealthy, English-speaking regions, so models describe objects and scenes from those regions accurately and others poorly — a wedding in one country is captioned as a wedding, and in another as "people in colourful clothing".
  • Attribute inference. Training captions frequently assert gender, age, and occupation from appearance, so models learn to do the same. The safety lesson treats this as the design problem it is.
  • Activity association. Captions associate activities with the demographics they most often co-occur with in the data, so a model can caption the same posture as "cooking" or as "working" depending on who is in the picture.

None of these is fixed by a bigger model. They are fixed, partially, by data auditing, stratified evaluation (see Safety, monitoring and follow-ups), and explicit output constraints.

Why evaluation needs multiple references

Consensus metrics are unstable with a single reference: your caption is compared against one person's phrasing, and a perfectly good caption that person did not write scores badly. Five references reduce that variance substantially and are why curated datasets provide them.

If you are building your own evaluation set, collect at least three references per image, from different annotators, with the instruction that they should not see each other's work.