Multimodal Vision-Language Models

Course Content

Multimodal Vision-Language Models

3 sections · 5 lessons

CLIP and Image-Text Alignment


You have been asked to build a photo filter for a marketplace app. Sellers upload pictures; you need to tag each one with a category so buyers can browse. There are 1,200 categories, from "vintage brass doorknob" to "toddler rain boots".

You do the obvious thing. You take a ResNet-50, replace the final layer with a 1,200-way classifier, gather 200 labelled photos per category — 240,000 images, three weeks of annotation contracts — and train. Top-1 accuracy is 84%. You ship it.

Six weeks later, product adds a category: "reusable silicone food wrap". Your model has 1,200 output slots. There is no 1,201st slot. To add one you must collect new labelled data, attach a new classification head, and retrain. That happens again the following month, and the month after.

Worse, a seller writes in asking why searching "something to keep sandwiches fresh" returns nothing. Your model does not know the phrase. It knows exactly 1,200 integers, and integer #847 happens to have the string "reusable silicone food wrap" glued to it in a lookup table somewhere. The model itself never saw that string. It has no idea what the words mean.

That is the wall every fixed-vocabulary vision model hits. CLIP is the design that walks around it.

One training batch: cosine similarity, image rows by caption columns0.310.04-0.020.080.060.280.03-0.01-0.030.050.340.020.090.010.040.3photo of a doorknobphoto of rain bootsphoto of a lampphoto of a scarfimage 1image 2image 3image 4Divide by the temperature, softmax across each row and down each column, cross-entropy against the diagonal.
The diagonal is pulled up and every off-diagonal cell pushed down, which is why batch size is the biggest lever: it sets how many negatives each pair is contrasted against.

The fixed-vocabulary problem, stated precisely

A standard image classifier ends with a linear layer that maps a feature vector to N scores, one per class. The class names are not part of the model. They live in a JSON file next to it. Internally the model has learned "the pattern that lights up output neuron 847", and that neuron is defined entirely by which training images had label 847.

Three consequences follow, and they are all painful:

  • The vocabulary is frozen at training time. Adding a class means changing the architecture and retraining.
  • The model cannot use language. "Golden retriever" and "labrador" are as unrelated to it as "golden retriever" and "traffic cone" — they are just different integers. Semantic closeness between labels is thrown away.
  • Every class needs its own labelled examples. Rare classes get few examples and poor accuracy, and there is no way to describe a class in words instead of demonstrating it in pictures.

The reframing: make the label a sentence

CLIP (Contrastive Language–Image Pre-training, released by OpenAI in 2021) changes the question. Instead of asking "which of my N slots does this image belong to?", it asks "how well does this image match this piece of text?"

Both the image and the text get turned into vectors in the same space. Matching becomes a dot product. And because the text side is a general-purpose language encoder, the "label" can be anything you can type — a word, a phrase, a full sentence, a category you invented five seconds ago.

A classifier learns a fixed set of buckets. CLIP learns a measuring device: a function that scores any image against any sentence. The buckets are something you supply at query time, not something baked into the weights.

Two towers, one shared space

CLIP is two separate neural networks that never talk to each other during a forward pass, plus one training objective that forces their outputs to line up.

Text
  image (224x224x3)              text ("a photo of a golden retriever")        |                                      |   Vision Transformer                    Text Transformer   (ViT-B/32: 49 patches                 (12 layers, 77-token    + CLS token, 12 layers)               context, causal masking)        |                                      |   pooled vector (768)                   EOS-token vector (512)        |                                      |   linear projection -> 512              linear projection -> 512        |                                      |   L2 normalise                          L2 normalise        \                                     /         \___________  512-dim  ____________/                    SHARED SPACE              similarity = dot product

The image encoder

Most CLIP variants use a Vision Transformer. ViT-B/32 chops a 224×224 image into 32×32 pixel patches — that gives 7×7 = 49 patches — flattens each patch, projects it to a token embedding, prepends a learnable [CLS] token, adds positional embeddings, and runs 12 transformer layers over the resulting 50 tokens. The final [CLS] vector is the image representation.

ViT-L/14 uses 14×14 patches, so 224/14 = 16 across, giving 256 patches — finer spatial detail at roughly 5× the compute. The 336-pixel variant pushes this to 24×24 = 576 patches.

The text encoder

A 12-layer transformer over byte-pair-encoded tokens, with a hard limit of 77 tokens including the start and end markers, so about 60 usable words. The representation is taken from the position of the end-of-text token. Text longer than 77 tokens is silently truncated by most libraries — a caption that gets cut in the middle contributes a mangled embedding and you will not get a warning.

Why L2 normalisation matters

Both projections are normalised to unit length before comparison, which turns the dot product into cosine similarity in [-1, 1]. Without it the model could cut its loss by inflating vector magnitudes rather than improving direction, and scores would be incomparable across queries. Normalisation forces meaning into direction only.

Contrastive training, with the arithmetic done

Here is the mechanism that does all the work, and it is worth going through slowly with real numbers, because almost every practical decision about CLIP — batch size, temperature, fine-tuning strategy — falls out of it.

Take a batch of N image–text pairs scraped from the web. CLIP was trained on 400 million such pairs. Encode all N images and all N texts. You now have an N×N matrix of cosine similarities. The N diagonal entries are the true pairs. The N² − N off-diagonal entries are in-batch negatives — free wrong answers, obtained without any extra labelling, simply because image i's caption is not image j's caption.

Take N = 4 for a batch we can compute by hand:

cosine simT1 "golden retriever"T2 "slice of pizza"T3 "Eiffel tower at night"T4 "tabby cat"
I1 dog photo0.300.100.050.22
I2 pizza photo0.080.280.060.09
I3 tower photo0.040.070.310.05
I4 cat photo0.240.090.060.27

Notice I1–T4 = 0.22 and I4–T1 = 0.24. Dogs and cats sit close together in this space. Those are the hard negatives — wrong answers that are nearly right.

Step 1: divide by the temperature

The raw similarities are squashed into a narrow band. CLIP divides every entry by a learned temperature τ, initialised at 0.07. Row 1 becomes:

0.300.07=4.286,0.100.07=1.429,0.050.07=0.714,0.220.07=3.143\frac{0.30}{0.07}=4.286,\quad \frac{0.10}{0.07}=1.429,\quad \frac{0.05}{0.07}=0.714,\quad \frac{0.22}{0.07}=3.143

Step 2: softmax across the row, then cross-entropy

Exponentiating: e^4.286 = 72.66, e^1.429 = 4.17, e^0.714 = 2.04, e^3.143 = 23.17. Sum = 102.04.

candidate text for image I1logitexpprobability
T1 golden retriever (correct)4.28672.660.712
T4 tabby cat (hard negative)3.14323.170.227
T2 slice of pizza1.4294.170.041
T3 Eiffel tower0.7142.040.020

Loss for this row = −ln(0.712) = 0.340. Doing the same for the other three rows gives 0.154 (pizza), 0.075 (tower) and 0.575 (cat). Mean image→text loss = 0.286.

Step 3: do it again down the columns

The matrix is not symmetric, so "which caption fits this image" and "which image fits this caption" are different questions with different competitors. Column T4 ("tabby cat") is contested by I1 (0.22) and won by I4 (0.27) — a much tighter race than row 4 saw. Column losses come out at 0.400, 0.176, 0.077 and 0.476, mean 0.282.

CLIP's loss is the average of both directions: (0.286 + 0.282) / 2 = 0.284.

For scale, a model that has learned nothing scores 1/N on every entry, giving loss ln(4) = 1.386. So 0.284 means this toy model has learned a great deal.

What the gradient actually does

For softmax cross-entropy the gradient on each logit is simply probability − target. For row 1 that means:

  • T1 (correct): 0.712 − 1 = −0.288 → push this similarity up
  • T4 (cat): +0.227 → push down, hard
  • T2 (pizza): +0.041 → push down, gently
  • T3 (tower): +0.020 → push down, barely

The cat absorbs 5.5× the corrective force of the pizza and 11× that of the tower. The softmax automatically routes the learning signal to whichever negatives are dangerous. Easy negatives contribute almost nothing — you spent compute encoding a pizza image and got a gradient of 0.02 out of it.

Contrastive learning is only as good as the hardest negative in the batch. Everything else is a rounding error.

Why the temperature is there at all

Run row 1 again with τ = 1 (no scaling). Logits are 0.30, 0.10, 0.05, 0.22; exponentials are 1.350, 1.105, 1.051, 1.246; sum 4.752; probability of the correct answer = 0.284. Loss = 1.259 — against a random baseline of 1.386. The model has the right answer ranked first and is still being told it is almost completely wrong. Gradients are small and roughly equal in every direction, so training crawls.

Now the other extreme, τ = 0.01. Logits become 30, 10, 5, 22. The correct answer's probability is 0.9997 and the loss is 0.0003. There is effectively no gradient — but the moment a single hard negative overtakes the correct pair, the loss explodes. Training becomes unstable. This is exactly why CLIP clamps its learned logit scale at 100, i.e. τ cannot go below 0.01.

temperature τP(correct) in row 1lossbehaviour
1.000.2841.259Near-random loss; flat, uninformative gradients
0.07 (CLIP init)0.7120.340Correct pair dominates but hard negative still pushes back
0.01 (clamp floor)0.99970.0003Gradient vanishes; brittle, spiky loss

τ is learned, not fixed — CLIP parameterises it as a log-scale value and lets gradient descent choose. It typically settles near 0.01, i.e. a logit scale close to 100.

Why batch size is the single biggest lever

CLIP was trained with a batch of 32,768 pairs. That number looks extravagant until you count what it buys.

batch size Nnegatives per anchorsimilarity entries per step (N²)random-guess loss ln(N)
25625565,5365.55
4,0964,09516,777,2168.32
32,76832,7671,073,741,82410.40

Going from 256 to 32,768 is 128× more encoder work per step but 16,384× more pairwise comparisons. Encoding cost grows linearly in N; the supervision signal grows quadratically. That is the whole trick.

The sharper argument is about hard negatives. Suppose that for any given caption, roughly 1 image in 1,000 in your corpus is a genuine hard negative:

  • N = 256: expected hard negatives per anchor = 255/1000 = 0.26. Three quarters of your training steps contain none at all, and as we saw, easy negatives produce gradients near 0.02.
  • N = 32,768: expected hard negatives per anchor = 32,767/1000 = 32.8. Every step teaches a fine distinction.

Hard negatives arrive by chance; a big batch buys enough lottery tickets. Without 256 GPUs the workarounds are gradient caching (encode in chunks, keep only the embeddings, then compute the full N×N loss), all-gathering embeddings across GPUs so the loss sees the global batch rather than a per-device shard, or a sigmoid loss as SigLIP uses — scoring each pair independently instead of normalising over the batch removes the batch-size dependence almost entirely.

Using a pretrained CLIP

Bash
pip install torch transformers pillow
Python
import torchfrom PIL import Imagefrom transformers import CLIPModel, CLIPProcessordevice = "cuda" if torch.cuda.is_available() else "cpu"model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32").to(device).eval()processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")image = Image.open("listing_4471.jpg").convert("RGB")labels = ["a photo of a brass doorknob",          "a photo of toddler rain boots",          "a photo of reusable silicone food wrap"]inputs = processor(text=labels, images=image, return_tensors="pt",                   padding=True, truncation=True).to(device)with torch.no_grad():    out = model(**inputs)# logits_per_image has shape (n_images, n_texts) and is already# cosine similarity multiplied by the learned logit scale (~100).probs = out.logits_per_image.softmax(dim=-1)for label, p in zip(labels, probs[0].tolist()):    print(f"{p:.3f}  {label}")

Three things about that output object trip people up:

what you getwhat it actually istrap
logits_per_imagecosine similarity × logit_scaleNot a probability. Divide by model.logit_scale.exp() to recover raw cosine.
.softmax(dim=-1)a distribution over the labels you suppliedIt always sums to 1, so it will confidently pick a label even for a blank white image.
image_embeds, text_embedsthe 512-dim projected vectorsAlready L2-normalised in transformers; normalising again is harmless, forgetting to when you build them yourself is not.

That second row is the most common production bug. If your label set is "cat / dog / bird" and someone uploads a car, CLIP returns a confident answer anyway. Add escape-hatch prompts ("a photo of something else", "a blank image") and threshold on the raw cosine similarity, not the softmax probability.

Which variant to use

checkpointpatchesembed dimapprox. ImageNet zero-shotuse when
clip-vit-base-patch3249512~63%High-throughput indexing, CPU inference, prototyping
clip-vit-base-patch16196512~68%Good default; noticeably better on small objects
clip-vit-large-patch14256768~75%Accuracy matters more than latency
clip-vit-large-patch14-336576768~76%Fine detail: text on packaging, small defects

Beyond OpenAI's originals, OpenCLIP (the open_clip library) ships models trained on LAION-2B and DataComp that beat them at matched size, and Google's SigLIP 2 (2025, e.g. google/siglip2-base-patch16-224, loaded with AutoModel) is usually the stronger retrieval choice today. Two SigLIP habits differ from CLIP: pad text with padding="max_length", as it was trained, and read each score as an independent sigmoid probability rather than a softmax over your labels. Whatever you pick, embeddings from different checkpoints are not interchangeable — pin the model version alongside your index.

Zero-shot classification and the wording problem

"Zero-shot" means the model classifies into categories it was never explicitly trained on, with zero labelled examples from you. The mechanism is exactly the code above: embed your candidate label sentences, embed the image, take the argmax.

The awkward part is that the wording of the label changes the answer. CLIP was trained on web alt-text, where captions read like sentences, not single dictionary words. The bare word "boots" puts it in a distribution it rarely saw.

The CLIP authors measured this on ImageNet. Wrapping labels in "A photo of a {label}." instead of the bare label added about 1.3 percentage points. Ensembling 80 templates — "a blurry photo of a {}", "art of a {}", "a photo of the large {}" — and averaging the resulting text embeddings added roughly 3.5 points more, close to 5 points total. That is a larger gain than several architecture upgrades, for free.

Python
import torch.nn.functional as FTEMPLATES = ["a photo of a {}.", "a blurry photo of a {}.",             "a close-up photo of a {}.", "a photo of the {}.",             "a cropped photo of a {}.", "a bright photo of a {}."]def class_embeddings(model, processor, classnames, device):    vecs = []    for name in classnames:        prompts = [t.format(name) for t in TEMPLATES]        toks = processor(text=prompts, return_tensors="pt",                         padding=True, truncation=True).to(device)        with torch.no_grad():            # transformers 5+: an output object; the projected vectors are .pooler_output            e = model.get_text_features(**toks).pooler_output        e = F.normalize(e, dim=-1)      # normalise each template        e = F.normalize(e.mean(dim=0), dim=-1)  # average, then re-normalise        vecs.append(e)    return torch.stack(vecs)            # (n_classes, dim)

Note the two normalisations. Averaging unit vectors produces something shorter than unit length, so you must re-normalise or the similarity scale drifts per class, systematically penalising classes whose templates disagree.

Failure modes worth knowing before you trust it

failurewhat happensmitigation
Typographic attackTape a paper reading "iPod" to an apple and CLIP calls it an iPod. It reads text in images and weights it heavily.Detect and mask rendered text, or add a text-in-image detector upstream
Bag-of-words behaviour"a horse riding an astronaut" and "an astronaut riding a horse" get near-identical embeddings; word order and relations are weakly encodedUse a generative VLM for anything relational; CLIP is for matching, not parsing
Counting"three cats" vs "five cats" barely differUse an object detector and count boxes
Forced choiceSoftmax over your labels always produces a winner, even for out-of-domain inputThreshold raw cosine similarity; include "none of these" prompts
Uncalibrated scoresA cosine of 0.28 may be excellent for one prompt set and mediocre for anotherCalibrate thresholds per prompt set on a held-out sample; never hardcode a global cut-off

Cross-modal search at scale

Because both modalities land in one space, search works in either direction with the same operation. The key efficiency point: text embeddings for a fixed label set are computed once, and image embeddings are computed once at upload time. Query time is a matrix multiply.

Python
import torch, torch.nn.functional as F@torch.no_grad()def index_images(model, processor, paths, device, batch_size=64):    """Encode a corpus once. Returns (n_images, dim) unit vectors."""    all_vecs = []    for i in range(0, len(paths), batch_size):        imgs = [Image.open(p).convert("RGB") for p in paths[i:i + batch_size]]        px = processor(images=imgs, return_tensors="pt").to(device)        with torch.autocast(device_type=device, dtype=torch.float16,                            enabled=(device == "cuda")):            v = model.get_image_features(**px).pooler_output        all_vecs.append(F.normalize(v.float(), dim=-1).cpu())    return torch.cat(all_vecs)@torch.no_grad()def search(model, processor, index, query, device, k=5):    toks = processor(text=[query], return_tensors="pt",                     padding=True, truncation=True).to(device)    q = F.normalize(model.get_text_features(**toks).pooler_output.float(), dim=-1).cpu()    scores = index @ q.T                 # (n_images, 1) cosine similarities    top = scores.squeeze(1).topk(k)    return list(zip(top.indices.tolist(), top.values.tolist()))

Some sizing arithmetic. One million images at 512 dimensions in float32 is 1,000,000 × 512 × 4 bytes = 2.05 GB; in float16, 1.02 GB, at almost no cost in retrieval quality. Brute-force search is a 1,000,000×512 by 512×1 matmul — about 512 million multiply-accumulates, single-digit milliseconds on a GPU. Brute force is genuinely fine up to a few million vectors; past roughly 10 million, move to an approximate index (HNSW or IVF-PQ) and accept 95–99% recall for sub-millisecond lookups.

Two batching rules pay for themselves immediately. Encoding images one at a time wastes most of a GPU; batch 64 with float16 autocast typically gives a 20–40× throughput improvement. And always wrap inference in torch.no_grad() — without it PyTorch builds an autograd graph per image and memory grows until the process dies.

Fine-tuning: when, and how not to wreck the model

CLIP is trained on web imagery: an enormous number of dogs, celebrities and landmarks, and comparatively few chest X-rays, circuit boards or satellite tiles. When your domain is narrow, zero-shot accuracy can be dismal even though the model is excellent in general.

Before fine-tuning, exhaust the cheaper options in order: better prompts, prompt ensembling, then a linear probe — freeze CLIP entirely, extract embeddings once, train logistic regression on top. A linear probe on a few hundred examples per class often lands within a couple of points of full fine-tuning, trains in seconds, and cannot damage the backbone.

If you do fine-tune the encoders, the failure mode is catastrophic forgetting: the model gets better at your 40 product categories and loses its general language grounding, so your open-ended search stops working. Guard against it:

knobsafe settingwhat goes wrong otherwise
Learning rate1e-6 to 5e-6 for encoders, 1e-4 for a new headAt 1e-4 on the backbone, the shared space collapses within a few hundred steps
What to unfreezeLast 2–4 transformer blocks and the projections; or LoRA adapters on attentionFull fine-tuning on a small dataset overfits and forgets
ObjectiveKeep the contrastive loss with text prompts, not a plain classifier headA classifier head throws away exactly the property you wanted — open vocabulary
Batch sizeAs large as fits; add explicit hard negatives if it is smallAt batch 16 you have 15 mostly-easy negatives and near-zero gradient
TemperatureFreeze it at the pretrained valueA learnable τ on a small dataset races to the clamp and destabilises training
MonitoringTrack a general-domain retrieval benchmark alongside your ownYou will not notice forgetting until search regressions reach users

If your fine-tuned CLIP scores brilliantly on your validation set and worse than zero-shot on any general query, you have not adapted the model — you have overwritten it.

What this changes about how you build

Go back to the marketplace. With CLIP the design inverts. Every uploaded photo is encoded once at upload time into a 512-float vector stored beside the listing — about 1 KB in float16. The category list is no longer part of the model; it is a config file of prompt strings. Product adds "reusable silicone food wrap" by adding one line and embedding one sentence. No retraining, no annotation contract, no deploy.

The search box changes too. "Something to keep sandwiches fresh" is just another sentence to embed, and it lands near the food-wrap images because the text encoder understands English, even though the image side never saw those words.

Two more shapes fall out of the same primitive:

  • Content moderation. Score each upload against policy prompts and benign prompts; route anything above a calibrated threshold to human review. New policy categories are new strings, so you respond to an emerging abuse pattern the same day, not the same quarter.
  • Visual recommendation. "More like this" becomes nearest-neighbour search in the image half of the space, and "more like this but in blue" becomes vector arithmetic: embed the source image, add a small multiple of the difference between "blue" and "red" text embeddings, re-normalise, search.

The discipline that keeps all of this working is boring but non-negotiable. Version the checkpoint alongside the index. Store raw cosine similarities, not softmax probabilities, so thresholds survive a change to the label set. Calibrate every threshold on held-out data rather than picking 0.3 because it is a round number. And keep a few hundred manually verified query/image pairs as a regression suite — it is the only thing that will tell you a model swap or prompt tweak made retrieval quietly worse.