Multimodal Vision-Language Models

Course Content

Multimodal Vision-Language Models

3 sections · 5 lessons

Captioning and Cross-Modal Retrieval


A stock photo library has 4.2 million images and a search box. The images were tagged years ago by whoever uploaded them, so a photograph of a woman laughing on a train platform is tagged woman, travel, lifestyle and nothing else.

A customer types "someone missing their train". Zero results. The words "missing" and "train" together do not appear in any tag on any image, and the one photo in the library that shows exactly that is invisible.

Two different fixes suggest themselves, and they are worth taking seriously as alternatives because most teams pick the wrong one first.

Fix A: generate a caption for every image and search the captions. Run a captioning model over all 4.2 million photos, store the sentences, and let the existing keyword search do its job. This is appealing because it reuses the search stack you already have.

Fix B: embed every image into a shared image-text vector space and search by vector similarity. No text is generated at all; the query sentence and the image land near each other directly.

Fix A produces "A woman standing on a train platform." That does not contain the word "missing" either — because the caption model described what was in the frame, not what was happening emotionally. Fix A has converted an image search problem into a text search problem and thrown away most of the image in the process.

Both techniques are useful. But they answer different questions, and confusing them is the most common design error in this area.

Searching 4.2 million badly tagged photographsQuery text, notags involvedText encoderto one vectorANN indexreturns top 100Cross-encoderreranks themTop 10 shownto the userStage one is fast and approximate; stage two is slow and accurate, and only ever sees 100 candidates.
Cross-modal retrieval skips captions entirely — the words never have to be written down for the photo to be findable.

What captioning is, and what it is not

Image captioning generates a natural-language description of an image. It is the model choosing what to say. That last part is what people underestimate: an image contains far more than any one sentence can carry, so every caption is an act of selection, and the model's training data decides what gets selected.

captioning is good atcaptioning is bad at
Alt-text for accessibilityBeing the sole index for search — it discards what it did not mention
Bulk metadata for un-tagged archivesPrecise counts, measurements and small text
Giving a human a quick sense of a fileAnything where an omission is a safety problem
Feeding a text-only downstream systemCapturing mood, intent or implied narrative

A caption is a lossy compression of an image into about fifteen words. Design as if the other 99% of the image is gone, because for anything reading only the caption, it is.

Two generations of architecture

The first generation, from roughly 2015, was CNN encoder plus RNN decoder. A convolutional network produced a feature grid; an LSTM generated words one at a time; an attention mechanism let each generated word look back at a weighted combination of grid cells, so "dog" could attend to the dog's pixels. It worked, and it was slow, brittle, and limited to the vocabulary and phrasings of its training captions.

The second generation is a vision transformer feeding a language model. The image becomes a set of patch vectors, those are projected into the language model's embedding space, and the language model generates text conditioned on them. The decoder is a general-purpose LLM, so it brings its full vocabulary, grammar and instruction-following.

CNN + LSTM with attentionViT + language model
VocabularyFixed, built from the caption corpusThe LLM's full tokeniser
Style controlNone — you get the training corpus's voicePrompt it: "one sentence", "for a 6-year-old", "SEO alt text"
GenerationSequential, no parallelism in the decoderSequential decode, but parallel prefill
Typical size~20M parameters1B–10B+ parameters
Handles unusual scenesPoorly — falls back to corpus clichésMuch better; still hallucinates

The practical consequence of the second row is bigger than it looks. Style is now a prompt, not a retraining job.

Python
import osimport torchfrom PIL import Imagefrom transformers import AutoProcessor, AutoModelForImageTextToTextMODEL_ID = os.environ.get("VLM_MODEL", "Qwen/Qwen3-VL-8B-Instruct")processor = AutoProcessor.from_pretrained(MODEL_ID)model = AutoModelForImageTextToText.from_pretrained(    MODEL_ID, dtype=torch.bfloat16, device_map="auto").eval()STYLES = {    "alt":      ("Write alt text for a screen reader. One sentence, under 125 "                 "characters. Describe only what is visible. Do not start with "                 "'an image of'.", 48, False),    "detailed": ("Describe this image in three sentences: the setting, the "                 "subjects, and the mood.", 160, True),    "keywords": ("List 8 search keywords for this image, comma separated, "                 "lowercase, no explanation.", 64, False),    "product":  ("Write a one-line product listing title: material, colour, "                 "item type. No marketing adjectives.", 40, False),}@torch.inference_mode()def caption(image: Image.Image, style: str = "alt") -> str:    instruction, max_tokens, creative = STYLES[style]    messages = [{"role": "user", "content": [        {"type": "image"}, {"type": "text", "text": instruction}]}]    prompt = processor.apply_chat_template(messages, add_generation_prompt=True)    inputs = processor(images=image, text=prompt,                       return_tensors="pt").to(model.device)    gen = dict(max_new_tokens=max_tokens)    gen.update(dict(do_sample=True, temperature=0.6, top_p=0.9)               if creative else dict(do_sample=False))    out = model.generate(**inputs, **gen)    cut = inputs["input_ids"].shape[-1]    return processor.decode(out[0][cut:], skip_special_tokens=True).strip()

Note that alt, keywords and product decode greedily while detailed samples. Anything a downstream system will parse, or anything where reproducibility matters, must be greedy. Sampling belongs only where a human reads the output and variety is a feature.

Dense captioning: one caption is not enough

A single caption for a busy street scene is nearly useless. Dense captioning produces a caption per region: run an object detector or region proposal network, then caption each crop.

Python
def dense_caption(image, boxes, pad=0.05):    """boxes: list of (x1, y1, x2, y2) in pixels, from a detector."""    w, h = image.size    out = []    for (x1, y1, x2, y2) in boxes:        px, py = int((x2 - x1) * pad), int((y2 - y1) * pad)        crop = image.crop((max(0, x1 - px), max(0, y1 - py),                           min(w, x2 + px), min(h, y2 + py)))        out.append({"box": (x1, y1, x2, y2), "caption": caption(crop, "alt")})    return out

The padding matters: a tight crop removes the context that makes an object interpretable, and a crop of a hand with no arm attached tends to be described as "a hand". The cost is linear in region count — 20 regions means 20 forward passes, so batch the crops and cap the region count by detector confidence. Dense captions are also far better retrieval material than a single global caption, because they preserve the small things a global caption drops.

Cross-modal retrieval: skip the words entirely

Retrieval uses a contrastive model such as CLIP, which encodes images and text into one shared vector space where matching pairs sit close together. No text is generated. The query sentence becomes a vector, the images are already vectors, and the search is a dot product.

Both directions use the same index:

text → imageimage → text
QueryA sentence the user typesAn uploaded photo
Index containsImage vectorsCaption / product-description vectors
Real useStock search, asset libraries, "find the photo with the red umbrella"Auto-tagging, "which of our 900 product descriptions matches this photo?"
Python
import torch, torch.nn.functional as Ffrom transformers import CLIPModel, CLIPProcessorclip = CLIPModel.from_pretrained("openai/clip-vit-base-patch32").eval()cproc = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")@torch.no_grad()def embed_images(images, batch_size=64):    vecs = []    for i in range(0, len(images), batch_size):        px = cproc(images=images[i:i + batch_size], return_tensors="pt")        # transformers 5+: the projected embedding is .pooler_output        vecs.append(F.normalize(clip.get_image_features(**px).pooler_output, dim=-1))    return torch.cat(vecs)                       # (N, 512), unit length@torch.no_grad()def embed_texts(texts):    tok = cproc(text=texts, return_tensors="pt", padding=True, truncation=True)    return F.normalize(clip.get_text_features(**tok).pooler_output, dim=-1)def topk(index, query_vec, k=10):    scores = (index @ query_vec.T).squeeze(1)    # cosine, since both unit norm    top = scores.topk(min(k, len(scores)))    return list(zip(top.indices.tolist(), top.values.tolist()))

Two traps live in those twenty lines. Forget the L2 normalisation and the dot product stops being cosine similarity — long vectors win regardless of direction, and your top result becomes whichever image happened to get a large-magnitude embedding. Mix embeddings from two different checkpoints in one index and the results are silently meaningless: nothing errors, the numbers look plausible, and the ranking is noise. Store the model identifier with the index and refuse to query on a mismatch.

Scaling past brute force

Brute force is a full scan. For 10 million images at 512 dimensions in float32:

  • Memory: 10,000,000 × 512 × 4 bytes = 20.5 GB
  • Per query: 10,000,000 × 512 = 5.12 billion multiply-accumulates, about 10.2 GFLOPs
  • At 100 queries per second that is over 1 TFLOP/s sustained, purely to search

Approximate nearest neighbour indexes trade a small amount of recall for orders of magnitude. The two that matter:

HNSW builds a navigable small-world graph in layers and walks it greedily. Query time is roughly logarithmic in the corpus size — microseconds rather than seconds — at very high recall. The cost is memory for the graph links: with 32 links per node stored as 4-byte ids in two directions, that is about 256 bytes per node, so 10 million nodes adds roughly 2.6 GB on top of the 20.5 GB of vectors.

IVF-PQ attacks memory instead. Cluster the vectors into nlist = 4,096 cells; at query time search only nprobe = 16 of them, which is 16/4,096 = 0.39% of the data — roughly a 256× reduction in comparisons. Then compress each vector by product quantisation: split 512 dimensions into 64 sub-vectors of 8 dimensions and replace each with a 1-byte codebook index. That takes each vector from 2,048 bytes to 64 bytes, a 32× compression, and the whole 10-million index from 20.5 GB to about 640 MB — small enough to sit in RAM on an ordinary server.

indexmemory (10M × 512)latencyrecall@10use when
Brute force (fp32)20.5 GB50–500 ms100%Under ~1M vectors, or exactness is required
Brute force (fp16)10.2 GB25–250 ms~100%Free 2× win; almost always worth doing
HNSW~23 GB<1 ms95–99%Latency matters and RAM is available
IVF-PQ~0.64 GB1–5 ms85–95%Corpus too big for RAM uncompressed

The failure mode here is tuning the index on a benchmark rather than on your own queries. Recall@10 of 95% sounds fine until you notice that the missing 5% is concentrated in rare, specific queries — precisely the ones where a user has something particular in mind and will notice its absence.

Two-stage retrieval: the pattern that actually ships

CLIP-style models are bi-encoders: image and text are encoded separately, so the whole corpus can be precomputed. That is what makes search fast. It is also what limits accuracy — the two encoders never see each other, so fine-grained interactions between specific words and specific regions are lost.

A cross-encoder — such as an image-text matching head that runs the pair through joint attention — is far more accurate and completely unusable as a first-stage search, because it must run once per candidate. Combine them:

Python
def search_two_stage(index, images, query, k=10, candidates=100):    q = embed_texts([query])    shortlist = topk(index, q, k=candidates)      # ~5 ms over 5M vectors    idxs = [i for i, _ in shortlist]    rescored = itm_score(query, [images[i] for i in idxs])  # cross-encoder    order = sorted(zip(idxs, rescored), key=lambda t: -t[1])    return order[:k]

The arithmetic justifies it. Suppose the cross-encoder scores 100 candidates in 200 ms. Running it over a 5-million-image corpus would take 5,000,000/100 × 200 ms = 10,000 seconds, or 2.8 hours per query. The two-stage version costs 5 ms plus 200 ms and recovers most of the accuracy. This is the standard architecture for any serious retrieval system.

Bi-encoders make search possible; cross-encoders make it good. Use the fast one to narrow the field and the accurate one to decide the winner.

Hybrid search: fusing vectors with keywords

Semantic search is bad at exact tokens. A query for part number SKU-4471-B is a string match, and a 512-dimensional embedding will happily return SKU-4471-D. Keyword search — BM25 — nails it. So run both and fuse the rankings.

Reciprocal rank fusion needs no score calibration at all, which is the point: dense cosine similarities and BM25 scores live on incomparable scales, so averaging them is meaningless. RRF averages ranks instead:

RRF(d)=∑i1k+ranki(d),k=60\text{RRF}(d) = \sum_{i} \frac{1}{k + \text{rank}_i(d)}, \qquad k = 60

Work an example. Document A ranks 1st in dense and 8th in keyword. Document B ranks 3rd in both.

  • A: 1/(60+1) + 1/(60+8) = 0.01639 + 0.01471 = 0.03110
  • B: 1/(60+3) + 1/(60+3) = 0.01587 + 0.01587 = 0.03175

B wins despite never placing first anywhere. That is the intended behaviour — agreement across two independent retrieval methods is stronger evidence than one method's enthusiasm. The constant k = 60 damps the top ranks so a single first place cannot dominate; lower it and RRF behaves more winner-take-all.

Captioning in other languages

The obvious approach — caption in English, then machine-translate — degrades in a specific way. "A woman in a sari at a wedding" translated into Hindi is grammatical and flat, because the English caption already collapsed the culturally specific detail that a Hindi speaker would have named directly.

approachhowtrade-off
Caption then translateEnglish VLM + translation modelCheapest; loses cultural specificity; errors compound across two models
Prompt a multilingual VLM directly"Describe this image in Hindi."Good on high-resource languages; quality falls off sharply on low-resource ones
Multilingual text encoder aligned to the image spaceDistil a multilingual encoder to match a frozen English CLIP text towerBest for retrieval: one image index serves every language

That third row deserves the emphasis for search. Because the image vectors never change, you align a multilingual text encoder to the existing space by training it to reproduce the English text encoder's embeddings for translated sentence pairs. Then a Spanish query and a Japanese query both land next to the same photograph, using one image index rather than one per language. For a 4.2-million-image library, that is the difference between 8.6 GB and 8.6 GB times however many languages you support.

Measuring captions without fooling yourself

Caption metrics are where teams most often mislead themselves, so work an example.

Reference: "a brown dog running across green grass" Candidate: "a dog runs across the grass"

Unigram precision: of the candidate's 6 tokens, a, dog, across, grass appear in the reference and runs, the do not. That is 4/6 = 0.667.

Bigram precision: the candidate's bigrams are a dog, dog runs, runs across, across the, the grass. The reference's are a brown, brown dog, dog running, running across, across green, green grass. Overlap: zero.

BLEU-4 multiplies precisions from 1 to 4 grams. With the bigram precision at 0, the product is 0, so BLEU-4 = 0 for a caption a human would call correct. The candidate's only sins were "runs" instead of "running" and dropping two adjectives.

metricwhat it measureswhere it breaks
BLEU-44-gram precision against referencesZero for correct paraphrases; needs many references to be stable
METEORUnigram match with stemming and synonymsHandles the example above; still surface-level
ROUGE-LLongest common subsequenceRecall-oriented; rewards long captions
CIDErTF-IDF weighted n-gram similarity across referencesThe caption standard; needs ~5 references per image
SPICEOverlap of parsed scene graphs (objects, attributes, relations)Closest to meaning; slow, and depends on a parser
CLIPScoreImage-caption cosine similarity, no reference neededInherits every CLIP blind spot — counts, word order, negation

CLIPScore is the one worth knowing because it needs no reference captions at all: score = 2.5 × max(cos(image, caption), 0). A cosine of 0.31 gives 2.5 × 0.31 = 0.775, which is in the range of a good caption. That means you can score captions on your own images without paying anyone to write ground truth. But it will happily give a high score to "three cats on a sofa" when there are two, because the underlying model cannot count.

Retrieval metrics are simpler and more honest. Recall@K asks whether the correct item appears in the top K — if the right image is in the top 5 for 830 of 1,000 test queries, Recall@5 = 83%. Mean reciprocal rank averages 1/rank: for ranks 1, 3, 2 and 10, MRR = (1 + 0.333 + 0.5 + 0.1)/4 = 0.483. Track both, because Recall@10 can look healthy while the correct answer sits stubbornly at position 8 and no real user ever scrolls that far.

Putting it together for the photo library

The right design uses both techniques for what each is good at, rather than choosing between them.

At ingest, each image is processed once. CLIP produces a 512-float vector — 1 KB in float16, so 4.3 GB for the whole 4.2-million library, comfortably in RAM. In parallel a VLM produces three artefacts: a screen-reader alt text, a keyword list, and dense captions for the top few detected regions. The vector goes into an HNSW index. The generated text goes into a BM25 index alongside the original human tags.

At query time, "someone missing their train" runs through both. The dense side finds the platform photograph because CLIP's text encoder places that sentence near images of people on platforms looking anxious. The keyword side finds anything whose caption or tags literally mention trains. RRF fuses the two rankings, a cross-encoder re-ranks the top 100, and the user sees ten results in about 250 ms.

The habits that keep it working are unglamorous. Version the embedding model alongside the index, because a checkpoint swap silently invalidates every stored vector. Keep captions as an additional signal rather than the primary index, since whatever the caption omitted is unfindable through it. Re-rank rather than trusting first-stage scores, because bi-encoder cosine values are not calibrated and a 0.31 for one query means something different from a 0.31 for another. And maintain a fixed set of a few hundred human-verified query-to-image pairs as a regression suite — it is the only mechanism that will catch a model upgrade or an index rebuild quietly making search worse, which it will, and which nobody will report until months later.