Course Content
Multimodal Vision-Language Models
3 sections · 5 lessons
Mini Project: Build a CLIP-Based Multimodal Image Assistant
A wedding photographer has 52,000 photographs on an external drive, organised into folders named by date. She needs one shot: the one where the groom's grandmother is laughing during the speeches. She knows it exists. She has no idea which of the 214 folders it is in, and the filenames are all IMG_4471.CR2.
Her current options are scrolling for two hours, or nothing. Every keyword search tool she has tried needs tags she never wrote.
You are going to build the thing that fixes this: a system where she types "an older woman laughing during a speech" and gets ten photographs in under a second, then asks a follow-up question about any one of them. Along the way you will hit every real engineering constraint in multimodal systems — the latency budget, the indexing throughput, the storage arithmetic, and the specific ways vector search goes wrong.
What you are building
Three capabilities, each backed by a different model behaviour:
| capability | what the user does | mechanism | when it runs |
|---|---|---|---|
| Semantic search | Types a sentence, gets ranked photos | Encode query text, cosine similarity against a precomputed image index | Query time — must be fast |
| Zero-shot tagging | Filters by "indoor", "black and white", "group shot" | Score each image against a list of label prompts, store the scores | Ingest time |
| Ask about a photo | Clicks a result and asks "how many people are at this table?" | Generative vision-language model on the single selected image | Query time, on demand only |
The first two use a contrastive model — two encoders that map images and sentences into one shared vector space so matching pairs sit close together. The third needs a generative model, which composes an answer rather than ranking options you supplied.
The architecture of this project is decided by one question: which work can be done once, at ingest, and which must happen while a user waits.
The latency budget, and what it forces
Set a target of 300 ms from keypress to results. Now cost out the parts on a modest GPU:
| operation | cost | can it be precomputed? |
|---|---|---|
| Encode 52,000 images with CLIP ViT-B/32 | ~90 s on a T4 at batch 64, fp16 | Yes — do it once at ingest |
| Encode one query sentence | ~8 ms | No, but it is cheap |
| Cosine similarity over 52,000 vectors | ~2 ms (52,000 × 512 = 26.6M MACs) | No, and it does not need to be |
| Load and serve 10 thumbnails | ~40 ms from local disk | Thumbnails, yes; generate at ingest |
| Generate one caption with a 7B VLM | ~900–1,500 ms | Yes — must be, or it blows the budget alone |
| Answer a question about one image | ~900–1,500 ms | No — it is on demand, so give it its own endpoint and its own spinner |
The budget adds up: 8 + 2 + 40 = 50 ms of unavoidable query-time work, leaving plenty of headroom. But the caption row alone is three to five times the whole 300 ms budget. That single number determines the architecture. Captioning and tagging move to ingest; only encoding, search and the on-demand question answering happen live.
The storage arithmetic
Also worth computing before you design a schema, because the intuition is usually backwards.
- Embeddings: 52,000 × 512 dims × 2 bytes (fp16) = 53 MB. Loads into RAM instantly.
- Thumbnails at 512px, ~40 KB each: 2.1 GB.
- Original RAW files: roughly 1.3 TB.
The vectors are the cheapest thing in the system by a factor of forty against the thumbnails alone. So do not compress them, do not reach for an approximate index, and do not build a vector database. At this scale a NumPy array on disk plus a brute-force matrix multiply is both faster and simpler than anything else, and it gives exact results. Approximate indexes start earning their keep somewhere above a few million vectors.
Choosing models
Pick these deliberately, because swapping the contrastive model later means re-indexing everything.
| role | suggested | why | when to change it |
|---|---|---|---|
| Contrastive encoder | openai/clip-vit-base-patch32 | 512-dim, fast enough to index 52k images in ~90 s, runs acceptably on CPU | Use patch16 or a SigLIP 2 model when retrieval quality matters more than index time |
| Generative VLM | A 7–9B open-weight chat VLM, e.g. Qwen/Qwen3-VL-8B-Instruct, read from config.py | Good quality-to-VRAM ratio; ~2 bytes per parameter in fp16/bf16 (~16 GB for 8B), about a quarter of that in 4-bit | Use a 2–4B model if you must run on 8 GB of VRAM |
| Index | NumPy float16 array | Exact, zero dependencies, 2 ms at this scale | Move to HNSW past ~2M vectors or if latency budget tightens |
The rule that will save you a day of confusion: the model identifier must be written into the index file, and search must refuse to run on a mismatch. Vectors from two different checkpoints occupy incompatible spaces. Nothing errors when you mix them — the numbers look plausible and the ranking is noise.
Project structure
photo-assistant/├── assistant/│ ├── __init__.py│ ├── encoder.py # CLIP wrapper: encode_images, encode_texts│ ├── index.py # build, save, load, search│ ├── ingest.py # walk folders, thumbnail, embed, tag, caption│ ├── vqa.py # generative VLM: ask(image, question)│ └── config.py # model ids, paths, label vocabulary, thresholds├── api/│ └── server.py # FastAPI: /search, /ask, /image/{id}├── ui/│ └── app.py # Gradio front end├── tests/│ ├── test_index.py│ └── eval_queries.json # 200 hand-verified query -> filename pairs├── data/ # index.npy, meta.sqlite, thumbs/├── Dockerfile└── requirements.txtThe separation that matters is assistant/ knowing nothing about HTTP and api/ knowing nothing about models. It means you can run the indexer as a batch job, test search without starting a server, and swap Gradio for anything else without touching the core.
Step 1 — Environment
1python -m venv .venv && source .venv/bin/activate2pip install torch torchvision transformers accelerate pillow \3 numpy fastapi "uvicorn[standard]" gradio pytest45python -c "import torch; print(torch.cuda.is_available(), torch.__version__)"If that prints False, everything still works — indexing 52,000 images on CPU takes roughly 35 minutes instead of 90 seconds, and the generative model will be slow enough that you should use a 2B checkpoint or skip that feature while developing.
Step 2 — The core
1# assistant/encoder.py2import torch, torch.nn.functional as F3from transformers import CLIPModel, CLIPProcessor45class Encoder:6 def __init__(self, model_id="openai/clip-vit-base-patch32", device=None):7 self.model_id = model_id8 self.device = device or ("cuda" if torch.cuda.is_available() else "cpu")9 self.model = CLIPModel.from_pretrained(model_id).to(self.device).eval()10 self.proc = CLIPProcessor.from_pretrained(model_id)1112 @torch.no_grad()13 def encode_images(self, pil_images, batch_size=64):14 out = []15 for i in range(0, len(pil_images), batch_size):16 px = self.proc(images=pil_images[i:i + batch_size],17 return_tensors="pt").to(self.device)18 # transformers 5+: the projected embedding is .pooler_output19 v = self.model.get_image_features(**px).pooler_output20 out.append(F.normalize(v, dim=-1).half().cpu())21 return torch.cat(out)2223 @torch.no_grad()24 def encode_texts(self, texts):25 tok = self.proc(text=texts, return_tensors="pt",26 padding=True, truncation=True).to(self.device)27 v = self.model.get_text_features(**tok).pooler_output28 return F.normalize(v, dim=-1).half().cpu()Every embedding is L2-normalised on the way out, so the dot product is cosine similarity and search is one matrix multiply. Skip that normalisation and long vectors win regardless of direction; your top result becomes whichever image happened to get a large-magnitude embedding, and the bug is invisible because the results are still plausible.
1# assistant/index.py2import json, numpy as np, torch34class Index:5 def __init__(self, vectors: np.ndarray, ids: list[str], model_id: str):6 assert vectors.shape[0] == len(ids)7 self.vectors, self.ids, self.model_id = vectors, ids, model_id89 def save(self, path_prefix: str):10 np.save(path_prefix + ".npy", self.vectors)11 with open(path_prefix + ".json", "w") as f:12 json.dump({"ids": self.ids, "model_id": self.model_id,13 "dim": int(self.vectors.shape[1])}, f)1415 @classmethod16 def load(cls, path_prefix: str, expected_model_id: str):17 vecs = np.load(path_prefix + ".npy")18 meta = json.load(open(path_prefix + ".json"))19 if meta["model_id"] != expected_model_id:20 raise RuntimeError(21 f"Index built with {meta['model_id']} but running "22 f"{expected_model_id}. Re-index before searching.")23 return cls(vecs, meta["ids"], meta["model_id"])2425 def search(self, query_vec: np.ndarray, k=10, min_score=0.0):26 scores = self.vectors @ query_vec.astype(self.vectors.dtype).ravel()27 k = min(k, scores.shape[0])28 top = np.argpartition(-scores, k - 1)[:k]29 top = top[np.argsort(-scores[top])]30 return [(self.ids[i], float(scores[i])) for i in top31 if float(scores[i]) >= min_score]Two details earn their place. The model_id check turns the silent-mismatch disaster into a loud startup error. And np.argpartition finds the top k in O(n) rather than sorting all 52,000 scores — small here, but it is free and it is the right habit.
Zero-shot tagging with prompt ensembling
1# assistant/ingest.py (excerpt)2LABELS = {3 "indoor": ["a photo taken indoors", "an indoor scene"],4 "outdoor": ["a photo taken outdoors", "an outdoor scene"],5 "group": ["a group photo of many people", "a crowd of people"],6 "portrait": ["a close-up portrait of one person"],7 "black_white": ["a black and white photograph", "a monochrome photo"],8 "dancing": ["people dancing", "a dance floor"],9 "speech": ["a person giving a speech", "someone making a toast"],10}1112def build_label_vectors(encoder):13 """Average several prompts per label, then re-normalise."""14 import torch.nn.functional as F15 names, vecs = [], []16 for name, prompts in LABELS.items():17 v = encoder.encode_texts(prompts).float() # (n_prompts, dim)18 vecs.append(F.normalize(v.mean(0), dim=-1))19 names.append(name)20 return names, torch.stack(vecs).numpy().astype("float16")The double normalisation is not redundant. Averaging unit vectors yields something shorter than unit length, so without the second normalize the similarity scale drifts per label — and labels whose prompts disagree with each other get systematically penalised, which shows up as a category that mysteriously never fires.
Prompt ensembling is worth the two extra lines. The CLIP authors measured roughly 5 percentage points of ImageNet zero-shot accuracy from prompt engineering and ensembling combined, which is larger than the gap between several model sizes.
The thresholding trap
You will want to write if score > 0.3: tag it. Do not hardcode that. Raw cosine similarities from a contrastive model are not calibrated — a score of 0.28 may be excellent for one prompt set and mediocre for another, and it shifts with the model checkpoint. Instead, hand-label 100 images per tag, sweep the threshold, and pick the point that hits your target precision:
1def choose_threshold(scores, labels, target_precision=0.90):2 """scores: model scores. labels: 1 if the tag truly applies."""3 order = np.argsort(-scores)4 tp = 05 best = 1.06 for rank, i in enumerate(order, start=1):7 tp += labels[i]8 if tp / rank >= target_precision:9 best = float(scores[i]) # lowest score still above target10 return bestStep 3 — The API
1# api/server.py2from fastapi import FastAPI, HTTPException3from fastapi.responses import FileResponse4from pydantic import BaseModel5from assistant.encoder import Encoder6from assistant.index import Index7from assistant import config, vqa89app = FastAPI()10encoder = Encoder(config.CLIP_MODEL_ID)11index = Index.load(config.INDEX_PREFIX, config.CLIP_MODEL_ID)1213class SearchRequest(BaseModel):14 query: str15 k: int = 1016 min_score: float = 0.201718@app.post("/search")19def search(req: SearchRequest):20 if not req.query.strip():21 raise HTTPException(400, "empty query")22 qv = encoder.encode_texts([req.query]).numpy()23 hits = index.search(qv, k=min(req.k, 50), min_score=req.min_score)24 return {"results": [{"id": i, "score": round(s, 4)} for i, s in hits]}2526class AskRequest(BaseModel):27 image_id: str28 question: str2930@app.post("/ask")31def ask(req: AskRequest):32 path = config.path_for(req.image_id)33 if path is None:34 raise HTTPException(404, "unknown image")35 return {"answer": vqa.ask(path, req.question)}3637@app.get("/image/{image_id}")38def image(image_id: str):39 path = config.thumb_for(image_id)40 if path is None:41 raise HTTPException(404, "unknown image")42 return FileResponse(path, media_type="image/jpeg")Note that /search and /ask are separate endpoints with wildly different latency profiles — 50 ms versus 1,200 ms. Merging them, so that every search also generated an answer, would make the fast path as slow as the slow path. Keep them apart and let the UI show a spinner only where one is warranted.
min_score defaults to a real value rather than zero because top-k always returns k results. Search for "a photograph of the surface of Mars" in a wedding archive and, with no floor, you will get ten confident-looking wedding photos. A floor lets the honest answer — nothing — come back.
Step 4 — The interface
1# ui/app.py2import gradio as gr, requests3API = "http://localhost:8000"45def do_search(query, k):6 r = requests.post(f"{API}/search", json={"query": query, "k": int(k)})7 hits = r.json()["results"]8 return [(f"{API}/image/{h['id']}", f"{h['score']:.3f}") for h in hits]910with gr.Blocks() as demo:11 gr.Markdown("### Photo assistant")12 with gr.Row():13 q = gr.Textbox(label="Describe the photo you want", scale=4)14 k = gr.Slider(4, 40, value=12, step=4, label="results")15 gallery = gr.Gallery(columns=4, height=560)16 q.submit(do_search, [q, k], gallery)1718 with gr.Row():19 img_id = gr.Textbox(label="Image id")20 question = gr.Textbox(label="Ask about this photo", scale=3)21 answer = gr.Textbox(label="Answer")22 question.submit(23 lambda i, t: requests.post(f"{API}/ask",24 json={"image_id": i, "question": t}).json()["answer"],25 [img_id, question], answer)2627demo.launch()Step 5 — Testing that means something
Unit tests on the index are necessary but easy. The test that actually tells you whether the system works is a retrieval regression suite: a fixed file of hand-verified query-to-image pairs, scored on every change.
1# tests/test_index.py2import json, numpy as np34def test_recall_at_k(index, encoder, k=5):5 cases = json.load(open("tests/eval_queries.json"))6 hits = 07 ranks = []8 for case in cases:9 qv = encoder.encode_texts([case["query"]]).numpy()10 got = [i for i, _ in index.search(qv, k=50)]11 if case["image_id"] in got[:k]:12 hits += 113 ranks.append(got.index(case["image_id"]) + 114 if case["image_id"] in got else 51)15 recall = hits / len(cases)16 mrr = sum(1.0 / r for r in ranks) / len(ranks)17 print(f"Recall@{k} = {recall:.3f} MRR = {mrr:.3f}")18 assert recall >= 0.75 # ratchet this up; never let it fallWork the numbers so you know what you are looking at. With 200 test queries and the correct image inside the top 5 for 166 of them, Recall@5 = 166/200 = 0.83. If four sample queries return the target at ranks 1, 3, 2 and 10, then MRR = (1 + 0.333 + 0.5 + 0.1)/4 = 0.483. Track both: Recall@10 can look healthy while the right answer sits at position 8, where nobody ever scrolls.
Without a regression suite you have no way to tell an improvement from a regression, and a vector search system degrades silently — nothing errors, results merely become slightly worse, and nobody reports it for months.
Failure modes you will actually hit
| symptom | cause | fix |
|---|---|---|
| Every query returns the same handful of photos | Embeddings not normalised, so magnitude dominates direction | L2-normalise on write; assert abs(norm − 1) < 1e-3 in the indexer |
| Results are plausible but never right | Index built with one checkpoint, queries encoded with another | The model_id check in Index.load |
| Nonsense queries return confident results | Top-k always returns k | A calibrated min_score floor |
| Indexing crashes at image 31,000 | A corrupt file, a CMYK JPEG, or an EXIF-rotated image | Image.open(p).convert("RGB") inside a try/except; log and skip |
| Photos come back sideways | EXIF orientation ignored, so CLIP saw a rotated image | PIL.ImageOps.exif_transpose before encoding and before thumbnailing |
Every .CR2 file is skipped | Pillow cannot decode camera RAW formats | Read the embedded JPEG preview or decode with rawpy, then encode and thumbnail that |
| Searching for "SKU-4471-B" fails | Semantic embeddings cannot do exact tokens | Add a keyword index and fuse rankings; this is what hybrid search is for |
| GPU out of memory after a few hundred images | Missing torch.no_grad(), so an autograd graph is retained per image | Decorate every inference function |
| Search latency creeps to seconds | Re-loading the model or the index per request | Load once at module import, as in the server above |
Deployment
1# Dockerfile2FROM python:3.11-slim3WORKDIR /app4RUN apt-get update && apt-get install -y --no-install-recommends \5 libgl1 libglib2.0-0 && rm -rf /var/lib/apt/lists/*6COPY requirements.txt .7RUN pip install --no-cache-dir -r requirements.txt8COPY . .9ENV HF_HOME=/models10EXPOSE 800011CMD ["uvicorn", "api.server:app", "--host", "0.0.0.0", "--port", "8000"]Two decisions here save real pain. Mount /models as a volume so the container does not re-download several gigabytes of weights on every restart — a cold start that downloads a 7B checkpoint takes minutes and will time out your orchestrator's health check. And keep data/ on a mounted volume too, so rebuilding the image does not throw away a 35-minute index.
Definition of done
| area | the bar |
|---|---|
| Search quality | Recall@5 at or above 0.75 on a 200-query hand-verified suite |
| Latency | p95 search under 300 ms; question answering has its own endpoint and its own expectations |
| Robustness | Indexing survives corrupt files, CMYK JPEGs and EXIF rotation without stopping |
| Honesty | An off-corpus query returns zero results, not ten confident wrong ones |
| Reproducibility | Model ids pinned; index refuses to load against a different checkpoint |
| Separation | Core logic importable and testable without starting a web server |
Once that holds, the interesting extensions are all cheap. Find similar photos is the same search with an image vector as the query instead of a text vector — a three-line change, since both live in the same space. Faceted filtering combines the stored zero-shot tag scores with the similarity ranking. Duplicate detection is a threshold on pairwise cosine similarity, which on a wedding archive of burst shots will find thousands. Multilingual search needs only a text encoder aligned to the same image space; the image index never changes.
What the project is really teaching
The models here are downloads. The engineering is everything else, and it generalises well beyond photographs.
You will have learned that the split between ingest-time and query-time work is the first architectural decision and it flows from arithmetic, not taste — a 1,200 ms caption in a 300 ms budget is not a tuning problem, it is a scheduling one. You will have learned that embeddings are only meaningful relative to the model that produced them, which is why the version check is load-bearing rather than defensive. You will have learned that unnormalised vectors and uncalibrated thresholds fail quietly, producing plausible results that are wrong, which is the most expensive kind of bug because nothing alerts you.
And you will have a regression suite of 200 verified query-image pairs. That artefact outlives the code. Every future change — a new checkpoint, a re-index, a prompt tweak, an approximate index swapped in when the archive reaches two million photos — gets measured against it. Systems built on vector similarity do not break loudly. They drift. The suite is the only thing that notices.