Course Content
Multimodal Vision-Language Models
3 sections · 5 lessons
CLIP and Image-Text Alignment
You have been asked to build a photo filter for a marketplace app. Sellers upload pictures; you need to tag each one with a category so buyers can browse. There are 1,200 categories, from "vintage brass doorknob" to "toddler rain boots".
You do the obvious thing. You take a ResNet-50, replace the final layer with a 1,200-way classifier, gather 200 labelled photos per category — 240,000 images, three weeks of annotation contracts — and train. Top-1 accuracy is 84%. You ship it.
Six weeks later, product adds a category: "reusable silicone food wrap". Your model has 1,200 output slots. There is no 1,201st slot. To add one you must collect new labelled data, attach a new classification head, and retrain. That happens again the following month, and the month after.
Worse, a seller writes in asking why searching "something to keep sandwiches fresh" returns nothing. Your model does not know the phrase. It knows exactly 1,200 integers, and integer #847 happens to have the string "reusable silicone food wrap" glued to it in a lookup table somewhere. The model itself never saw that string. It has no idea what the words mean.
That is the wall every fixed-vocabulary vision model hits. CLIP is the design that walks around it.
The fixed-vocabulary problem, stated precisely
A standard image classifier ends with a linear layer that maps a feature vector to N scores, one per class. The class names are not part of the model. They live in a JSON file next to it. Internally the model has learned "the pattern that lights up output neuron 847", and that neuron is defined entirely by which training images had label 847.
Three consequences follow, and they are all painful:
- The vocabulary is frozen at training time. Adding a class means changing the architecture and retraining.
- The model cannot use language. "Golden retriever" and "labrador" are as unrelated to it as "golden retriever" and "traffic cone" — they are just different integers. Semantic closeness between labels is thrown away.
- Every class needs its own labelled examples. Rare classes get few examples and poor accuracy, and there is no way to describe a class in words instead of demonstrating it in pictures.
The reframing: make the label a sentence
CLIP (Contrastive Language–Image Pre-training, released by OpenAI in 2021) changes the question. Instead of asking "which of my N slots does this image belong to?", it asks "how well does this image match this piece of text?"
Both the image and the text get turned into vectors in the same space. Matching becomes a dot product. And because the text side is a general-purpose language encoder, the "label" can be anything you can type — a word, a phrase, a full sentence, a category you invented five seconds ago.
A classifier learns a fixed set of buckets. CLIP learns a measuring device: a function that scores any image against any sentence. The buckets are something you supply at query time, not something baked into the weights.
Two towers, one shared space
CLIP is two separate neural networks that never talk to each other during a forward pass, plus one training objective that forces their outputs to line up.
image (224x224x3) text ("a photo of a golden retriever") | | Vision Transformer Text Transformer (ViT-B/32: 49 patches (12 layers, 77-token + CLS token, 12 layers) context, causal masking) | | pooled vector (768) EOS-token vector (512) | | linear projection -> 512 linear projection -> 512 | | L2 normalise L2 normalise \ / \___________ 512-dim ____________/ SHARED SPACE similarity = dot productThe image encoder
Most CLIP variants use a Vision Transformer. ViT-B/32 chops a 224×224 image into 32×32 pixel patches — that gives 7×7 = 49 patches — flattens each patch, projects it to a token embedding, prepends a learnable [CLS] token, adds positional embeddings, and runs 12 transformer layers over the resulting 50 tokens. The final [CLS] vector is the image representation.
ViT-L/14 uses 14×14 patches, so 224/14 = 16 across, giving 256 patches — finer spatial detail at roughly 5× the compute. The 336-pixel variant pushes this to 24×24 = 576 patches.
The text encoder
A 12-layer transformer over byte-pair-encoded tokens, with a hard limit of 77 tokens including the start and end markers, so about 60 usable words. The representation is taken from the position of the end-of-text token. Text longer than 77 tokens is silently truncated by most libraries — a caption that gets cut in the middle contributes a mangled embedding and you will not get a warning.
Why L2 normalisation matters
Both projections are normalised to unit length before comparison, which turns the dot product into cosine similarity in [-1, 1]. Without it the model could cut its loss by inflating vector magnitudes rather than improving direction, and scores would be incomparable across queries. Normalisation forces meaning into direction only.
Contrastive training, with the arithmetic done
Here is the mechanism that does all the work, and it is worth going through slowly with real numbers, because almost every practical decision about CLIP — batch size, temperature, fine-tuning strategy — falls out of it.
Take a batch of N image–text pairs scraped from the web. CLIP was trained on 400 million such pairs. Encode all N images and all N texts. You now have an N×N matrix of cosine similarities. The N diagonal entries are the true pairs. The N² − N off-diagonal entries are in-batch negatives — free wrong answers, obtained without any extra labelling, simply because image i's caption is not image j's caption.
Take N = 4 for a batch we can compute by hand:
| cosine sim | T1 "golden retriever" | T2 "slice of pizza" | T3 "Eiffel tower at night" | T4 "tabby cat" |
|---|---|---|---|---|
| I1 dog photo | 0.30 | 0.10 | 0.05 | 0.22 |
| I2 pizza photo | 0.08 | 0.28 | 0.06 | 0.09 |
| I3 tower photo | 0.04 | 0.07 | 0.31 | 0.05 |
| I4 cat photo | 0.24 | 0.09 | 0.06 | 0.27 |
Notice I1–T4 = 0.22 and I4–T1 = 0.24. Dogs and cats sit close together in this space. Those are the hard negatives — wrong answers that are nearly right.
Step 1: divide by the temperature
The raw similarities are squashed into a narrow band. CLIP divides every entry by a learned temperature τ, initialised at 0.07. Row 1 becomes:
Step 2: softmax across the row, then cross-entropy
Exponentiating: e^4.286 = 72.66, e^1.429 = 4.17, e^0.714 = 2.04, e^3.143 = 23.17. Sum = 102.04.
| candidate text for image I1 | logit | exp | probability |
|---|---|---|---|
| T1 golden retriever (correct) | 4.286 | 72.66 | 0.712 |
| T4 tabby cat (hard negative) | 3.143 | 23.17 | 0.227 |
| T2 slice of pizza | 1.429 | 4.17 | 0.041 |
| T3 Eiffel tower | 0.714 | 2.04 | 0.020 |
Loss for this row = −ln(0.712) = 0.340. Doing the same for the other three rows gives 0.154 (pizza), 0.075 (tower) and 0.575 (cat). Mean image→text loss = 0.286.
Step 3: do it again down the columns
The matrix is not symmetric, so "which caption fits this image" and "which image fits this caption" are different questions with different competitors. Column T4 ("tabby cat") is contested by I1 (0.22) and won by I4 (0.27) — a much tighter race than row 4 saw. Column losses come out at 0.400, 0.176, 0.077 and 0.476, mean 0.282.
CLIP's loss is the average of both directions: (0.286 + 0.282) / 2 = 0.284.
For scale, a model that has learned nothing scores 1/N on every entry, giving loss ln(4) = 1.386. So 0.284 means this toy model has learned a great deal.
What the gradient actually does
For softmax cross-entropy the gradient on each logit is simply probability − target. For row 1 that means:
- T1 (correct): 0.712 − 1 = −0.288 → push this similarity up
- T4 (cat): +0.227 → push down, hard
- T2 (pizza): +0.041 → push down, gently
- T3 (tower): +0.020 → push down, barely
The cat absorbs 5.5× the corrective force of the pizza and 11× that of the tower. The softmax automatically routes the learning signal to whichever negatives are dangerous. Easy negatives contribute almost nothing — you spent compute encoding a pizza image and got a gradient of 0.02 out of it.
Contrastive learning is only as good as the hardest negative in the batch. Everything else is a rounding error.
Why the temperature is there at all
Run row 1 again with τ = 1 (no scaling). Logits are 0.30, 0.10, 0.05, 0.22; exponentials are 1.350, 1.105, 1.051, 1.246; sum 4.752; probability of the correct answer = 0.284. Loss = 1.259 — against a random baseline of 1.386. The model has the right answer ranked first and is still being told it is almost completely wrong. Gradients are small and roughly equal in every direction, so training crawls.
Now the other extreme, τ = 0.01. Logits become 30, 10, 5, 22. The correct answer's probability is 0.9997 and the loss is 0.0003. There is effectively no gradient — but the moment a single hard negative overtakes the correct pair, the loss explodes. Training becomes unstable. This is exactly why CLIP clamps its learned logit scale at 100, i.e. τ cannot go below 0.01.
| temperature τ | P(correct) in row 1 | loss | behaviour |
|---|---|---|---|
| 1.00 | 0.284 | 1.259 | Near-random loss; flat, uninformative gradients |
| 0.07 (CLIP init) | 0.712 | 0.340 | Correct pair dominates but hard negative still pushes back |
| 0.01 (clamp floor) | 0.9997 | 0.0003 | Gradient vanishes; brittle, spiky loss |
τ is learned, not fixed — CLIP parameterises it as a log-scale value and lets gradient descent choose. It typically settles near 0.01, i.e. a logit scale close to 100.
Why batch size is the single biggest lever
CLIP was trained with a batch of 32,768 pairs. That number looks extravagant until you count what it buys.
| batch size N | negatives per anchor | similarity entries per step (N²) | random-guess loss ln(N) |
|---|---|---|---|
| 256 | 255 | 65,536 | 5.55 |
| 4,096 | 4,095 | 16,777,216 | 8.32 |
| 32,768 | 32,767 | 1,073,741,824 | 10.40 |
Going from 256 to 32,768 is 128× more encoder work per step but 16,384× more pairwise comparisons. Encoding cost grows linearly in N; the supervision signal grows quadratically. That is the whole trick.
The sharper argument is about hard negatives. Suppose that for any given caption, roughly 1 image in 1,000 in your corpus is a genuine hard negative:
- N = 256: expected hard negatives per anchor = 255/1000 = 0.26. Three quarters of your training steps contain none at all, and as we saw, easy negatives produce gradients near 0.02.
- N = 32,768: expected hard negatives per anchor = 32,767/1000 = 32.8. Every step teaches a fine distinction.
Hard negatives arrive by chance; a big batch buys enough lottery tickets. Without 256 GPUs the workarounds are gradient caching (encode in chunks, keep only the embeddings, then compute the full N×N loss), all-gathering embeddings across GPUs so the loss sees the global batch rather than a per-device shard, or a sigmoid loss as SigLIP uses — scoring each pair independently instead of normalising over the batch removes the batch-size dependence almost entirely.
Using a pretrained CLIP
pip install torch transformers pillow1import torch2from PIL import Image3from transformers import CLIPModel, CLIPProcessor45device = "cuda" if torch.cuda.is_available() else "cpu"6model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32").to(device).eval()7processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")89image = Image.open("listing_4471.jpg").convert("RGB")10labels = ["a photo of a brass doorknob",11 "a photo of toddler rain boots",12 "a photo of reusable silicone food wrap"]1314inputs = processor(text=labels, images=image, return_tensors="pt",15 padding=True, truncation=True).to(device)1617with torch.no_grad():18 out = model(**inputs)1920# logits_per_image has shape (n_images, n_texts) and is already21# cosine similarity multiplied by the learned logit scale (~100).22probs = out.logits_per_image.softmax(dim=-1)23for label, p in zip(labels, probs[0].tolist()):24 print(f"{p:.3f} {label}")Three things about that output object trip people up:
| what you get | what it actually is | trap |
|---|---|---|
logits_per_image | cosine similarity × logit_scale | Not a probability. Divide by model.logit_scale.exp() to recover raw cosine. |
.softmax(dim=-1) | a distribution over the labels you supplied | It always sums to 1, so it will confidently pick a label even for a blank white image. |
image_embeds, text_embeds | the 512-dim projected vectors | Already L2-normalised in transformers; normalising again is harmless, forgetting to when you build them yourself is not. |
That second row is the most common production bug. If your label set is "cat / dog / bird" and someone uploads a car, CLIP returns a confident answer anyway. Add escape-hatch prompts ("a photo of something else", "a blank image") and threshold on the raw cosine similarity, not the softmax probability.
Which variant to use
| checkpoint | patches | embed dim | approx. ImageNet zero-shot | use when |
|---|---|---|---|---|
clip-vit-base-patch32 | 49 | 512 | ~63% | High-throughput indexing, CPU inference, prototyping |
clip-vit-base-patch16 | 196 | 512 | ~68% | Good default; noticeably better on small objects |
clip-vit-large-patch14 | 256 | 768 | ~75% | Accuracy matters more than latency |
clip-vit-large-patch14-336 | 576 | 768 | ~76% | Fine detail: text on packaging, small defects |
Beyond OpenAI's originals, OpenCLIP (the open_clip library) ships models trained on LAION-2B and DataComp that beat them at matched size, and Google's SigLIP 2 (2025, e.g. google/siglip2-base-patch16-224, loaded with AutoModel) is usually the stronger retrieval choice today. Two SigLIP habits differ from CLIP: pad text with padding="max_length", as it was trained, and read each score as an independent sigmoid probability rather than a softmax over your labels. Whatever you pick, embeddings from different checkpoints are not interchangeable — pin the model version alongside your index.
Zero-shot classification and the wording problem
"Zero-shot" means the model classifies into categories it was never explicitly trained on, with zero labelled examples from you. The mechanism is exactly the code above: embed your candidate label sentences, embed the image, take the argmax.
The awkward part is that the wording of the label changes the answer. CLIP was trained on web alt-text, where captions read like sentences, not single dictionary words. The bare word "boots" puts it in a distribution it rarely saw.
The CLIP authors measured this on ImageNet. Wrapping labels in "A photo of a {label}." instead of the bare label added about 1.3 percentage points. Ensembling 80 templates — "a blurry photo of a {}", "art of a {}", "a photo of the large {}" — and averaging the resulting text embeddings added roughly 3.5 points more, close to 5 points total. That is a larger gain than several architecture upgrades, for free.
1import torch.nn.functional as F23TEMPLATES = ["a photo of a {}.", "a blurry photo of a {}.",4 "a close-up photo of a {}.", "a photo of the {}.",5 "a cropped photo of a {}.", "a bright photo of a {}."]67def class_embeddings(model, processor, classnames, device):8 vecs = []9 for name in classnames:10 prompts = [t.format(name) for t in TEMPLATES]11 toks = processor(text=prompts, return_tensors="pt",12 padding=True, truncation=True).to(device)13 with torch.no_grad():14 # transformers 5+: an output object; the projected vectors are .pooler_output15 e = model.get_text_features(**toks).pooler_output16 e = F.normalize(e, dim=-1) # normalise each template17 e = F.normalize(e.mean(dim=0), dim=-1) # average, then re-normalise18 vecs.append(e)19 return torch.stack(vecs) # (n_classes, dim)Note the two normalisations. Averaging unit vectors produces something shorter than unit length, so you must re-normalise or the similarity scale drifts per class, systematically penalising classes whose templates disagree.
Failure modes worth knowing before you trust it
| failure | what happens | mitigation |
|---|---|---|
| Typographic attack | Tape a paper reading "iPod" to an apple and CLIP calls it an iPod. It reads text in images and weights it heavily. | Detect and mask rendered text, or add a text-in-image detector upstream |
| Bag-of-words behaviour | "a horse riding an astronaut" and "an astronaut riding a horse" get near-identical embeddings; word order and relations are weakly encoded | Use a generative VLM for anything relational; CLIP is for matching, not parsing |
| Counting | "three cats" vs "five cats" barely differ | Use an object detector and count boxes |
| Forced choice | Softmax over your labels always produces a winner, even for out-of-domain input | Threshold raw cosine similarity; include "none of these" prompts |
| Uncalibrated scores | A cosine of 0.28 may be excellent for one prompt set and mediocre for another | Calibrate thresholds per prompt set on a held-out sample; never hardcode a global cut-off |
Cross-modal search at scale
Because both modalities land in one space, search works in either direction with the same operation. The key efficiency point: text embeddings for a fixed label set are computed once, and image embeddings are computed once at upload time. Query time is a matrix multiply.
1import torch, torch.nn.functional as F23@torch.no_grad()4def index_images(model, processor, paths, device, batch_size=64):5 """Encode a corpus once. Returns (n_images, dim) unit vectors."""6 all_vecs = []7 for i in range(0, len(paths), batch_size):8 imgs = [Image.open(p).convert("RGB") for p in paths[i:i + batch_size]]9 px = processor(images=imgs, return_tensors="pt").to(device)10 with torch.autocast(device_type=device, dtype=torch.float16,11 enabled=(device == "cuda")):12 v = model.get_image_features(**px).pooler_output13 all_vecs.append(F.normalize(v.float(), dim=-1).cpu())14 return torch.cat(all_vecs)1516@torch.no_grad()17def search(model, processor, index, query, device, k=5):18 toks = processor(text=[query], return_tensors="pt",19 padding=True, truncation=True).to(device)20 q = F.normalize(model.get_text_features(**toks).pooler_output.float(), dim=-1).cpu()21 scores = index @ q.T # (n_images, 1) cosine similarities22 top = scores.squeeze(1).topk(k)23 return list(zip(top.indices.tolist(), top.values.tolist()))Some sizing arithmetic. One million images at 512 dimensions in float32 is 1,000,000 × 512 × 4 bytes = 2.05 GB; in float16, 1.02 GB, at almost no cost in retrieval quality. Brute-force search is a 1,000,000×512 by 512×1 matmul — about 512 million multiply-accumulates, single-digit milliseconds on a GPU. Brute force is genuinely fine up to a few million vectors; past roughly 10 million, move to an approximate index (HNSW or IVF-PQ) and accept 95–99% recall for sub-millisecond lookups.
Two batching rules pay for themselves immediately. Encoding images one at a time wastes most of a GPU; batch 64 with float16 autocast typically gives a 20–40× throughput improvement. And always wrap inference in torch.no_grad() — without it PyTorch builds an autograd graph per image and memory grows until the process dies.
Fine-tuning: when, and how not to wreck the model
CLIP is trained on web imagery: an enormous number of dogs, celebrities and landmarks, and comparatively few chest X-rays, circuit boards or satellite tiles. When your domain is narrow, zero-shot accuracy can be dismal even though the model is excellent in general.
Before fine-tuning, exhaust the cheaper options in order: better prompts, prompt ensembling, then a linear probe — freeze CLIP entirely, extract embeddings once, train logistic regression on top. A linear probe on a few hundred examples per class often lands within a couple of points of full fine-tuning, trains in seconds, and cannot damage the backbone.
If you do fine-tune the encoders, the failure mode is catastrophic forgetting: the model gets better at your 40 product categories and loses its general language grounding, so your open-ended search stops working. Guard against it:
| knob | safe setting | what goes wrong otherwise |
|---|---|---|
| Learning rate | 1e-6 to 5e-6 for encoders, 1e-4 for a new head | At 1e-4 on the backbone, the shared space collapses within a few hundred steps |
| What to unfreeze | Last 2–4 transformer blocks and the projections; or LoRA adapters on attention | Full fine-tuning on a small dataset overfits and forgets |
| Objective | Keep the contrastive loss with text prompts, not a plain classifier head | A classifier head throws away exactly the property you wanted — open vocabulary |
| Batch size | As large as fits; add explicit hard negatives if it is small | At batch 16 you have 15 mostly-easy negatives and near-zero gradient |
| Temperature | Freeze it at the pretrained value | A learnable τ on a small dataset races to the clamp and destabilises training |
| Monitoring | Track a general-domain retrieval benchmark alongside your own | You will not notice forgetting until search regressions reach users |
If your fine-tuned CLIP scores brilliantly on your validation set and worse than zero-shot on any general query, you have not adapted the model — you have overwritten it.
What this changes about how you build
Go back to the marketplace. With CLIP the design inverts. Every uploaded photo is encoded once at upload time into a 512-float vector stored beside the listing — about 1 KB in float16. The category list is no longer part of the model; it is a config file of prompt strings. Product adds "reusable silicone food wrap" by adding one line and embedding one sentence. No retraining, no annotation contract, no deploy.
The search box changes too. "Something to keep sandwiches fresh" is just another sentence to embed, and it lands near the food-wrap images because the text encoder understands English, even though the image side never saw those words.
Two more shapes fall out of the same primitive:
- Content moderation. Score each upload against policy prompts and benign prompts; route anything above a calibrated threshold to human review. New policy categories are new strings, so you respond to an emerging abuse pattern the same day, not the same quarter.
- Visual recommendation. "More like this" becomes nearest-neighbour search in the image half of the space, and "more like this but in blue" becomes vector arithmetic: embed the source image, add a small multiple of the difference between
"blue"and"red"text embeddings, re-normalise, search.
The discipline that keeps all of this working is boring but non-negotiable. Version the checkpoint alongside the index. Store raw cosine similarities, not softmax probabilities, so thresholds survive a change to the label set. Calibrate every threshold on held-out data rather than picking 0.3 because it is a round number. And keep a few hundred manually verified query/image pairs as a regression suite — it is the only thing that will tell you a model swap or prompt tweak made retrieval quietly worse.