Multimodal Vision-Language Models

Course Content

Multimodal Vision-Language Models

3 sections · 5 lessons

Mini Project: Build a CLIP-Based Multimodal Image Assistant


A wedding photographer has 52,000 photographs on an external drive, organised into folders named by date. She needs one shot: the one where the groom's grandmother is laughing during the speeches. She knows it exists. She has no idea which of the 214 folders it is in, and the filenames are all IMG_4471.CR2.

Her current options are scrolling for two hours, or nothing. Every keyword search tool she has tried needs tags she never wrote.

You are going to build the thing that fixes this: a system where she types "an older woman laughing during a speech" and gets ten photographs in under a second, then asks a follow-up question about any one of them. Along the way you will hit every real engineering constraint in multimodal systems — the latency budget, the indexing throughput, the storage arithmetic, and the specific ways vector search goes wrong.

The storage arithmetic for 52,000 wedding photos52,000images on a drive512 floatsper embeddingfp16: about 53 MBThumbnails: 2.1GB, forty times moreFits in RAM —no vector DBtopbottomEmbedding is the one-off cost; a query is one text encode plus a dot product over the whole matrix.
At this scale brute-force cosine over an in-memory matrix beats any index, and the hard part is choosing a similarity threshold that is not arbitrary.

What you are building

Three capabilities, each backed by a different model behaviour:

capabilitywhat the user doesmechanismwhen it runs
Semantic searchTypes a sentence, gets ranked photosEncode query text, cosine similarity against a precomputed image indexQuery time — must be fast
Zero-shot taggingFilters by "indoor", "black and white", "group shot"Score each image against a list of label prompts, store the scoresIngest time
Ask about a photoClicks a result and asks "how many people are at this table?"Generative vision-language model on the single selected imageQuery time, on demand only

The first two use a contrastive model — two encoders that map images and sentences into one shared vector space so matching pairs sit close together. The third needs a generative model, which composes an answer rather than ranking options you supplied.

The architecture of this project is decided by one question: which work can be done once, at ingest, and which must happen while a user waits.

The latency budget, and what it forces

Set a target of 300 ms from keypress to results. Now cost out the parts on a modest GPU:

operationcostcan it be precomputed?
Encode 52,000 images with CLIP ViT-B/32~90 s on a T4 at batch 64, fp16Yes — do it once at ingest
Encode one query sentence~8 msNo, but it is cheap
Cosine similarity over 52,000 vectors~2 ms (52,000 × 512 = 26.6M MACs)No, and it does not need to be
Load and serve 10 thumbnails~40 ms from local diskThumbnails, yes; generate at ingest
Generate one caption with a 7B VLM~900–1,500 msYes — must be, or it blows the budget alone
Answer a question about one image~900–1,500 msNo — it is on demand, so give it its own endpoint and its own spinner

The budget adds up: 8 + 2 + 40 = 50 ms of unavoidable query-time work, leaving plenty of headroom. But the caption row alone is three to five times the whole 300 ms budget. That single number determines the architecture. Captioning and tagging move to ingest; only encoding, search and the on-demand question answering happen live.

The storage arithmetic

Also worth computing before you design a schema, because the intuition is usually backwards.

  • Embeddings: 52,000 × 512 dims × 2 bytes (fp16) = 53 MB. Loads into RAM instantly.
  • Thumbnails at 512px, ~40 KB each: 2.1 GB.
  • Original RAW files: roughly 1.3 TB.

The vectors are the cheapest thing in the system by a factor of forty against the thumbnails alone. So do not compress them, do not reach for an approximate index, and do not build a vector database. At this scale a NumPy array on disk plus a brute-force matrix multiply is both faster and simpler than anything else, and it gives exact results. Approximate indexes start earning their keep somewhere above a few million vectors.

Choosing models

Pick these deliberately, because swapping the contrastive model later means re-indexing everything.

rolesuggestedwhywhen to change it
Contrastive encoderopenai/clip-vit-base-patch32512-dim, fast enough to index 52k images in ~90 s, runs acceptably on CPUUse patch16 or a SigLIP 2 model when retrieval quality matters more than index time
Generative VLMA 7–9B open-weight chat VLM, e.g. Qwen/Qwen3-VL-8B-Instruct, read from config.pyGood quality-to-VRAM ratio; ~2 bytes per parameter in fp16/bf16 (~16 GB for 8B), about a quarter of that in 4-bitUse a 2–4B model if you must run on 8 GB of VRAM
IndexNumPy float16 arrayExact, zero dependencies, 2 ms at this scaleMove to HNSW past ~2M vectors or if latency budget tightens

The rule that will save you a day of confusion: the model identifier must be written into the index file, and search must refuse to run on a mismatch. Vectors from two different checkpoints occupy incompatible spaces. Nothing errors when you mix them — the numbers look plausible and the ranking is noise.

Project structure

Text
photo-assistant/├── assistant/│   ├── __init__.py│   ├── encoder.py       # CLIP wrapper: encode_images, encode_texts│   ├── index.py         # build, save, load, search│   ├── ingest.py        # walk folders, thumbnail, embed, tag, caption│   ├── vqa.py           # generative VLM: ask(image, question)│   └── config.py        # model ids, paths, label vocabulary, thresholds├── api/│   └── server.py        # FastAPI: /search, /ask, /image/{id}├── ui/│   └── app.py           # Gradio front end├── tests/│   ├── test_index.py│   └── eval_queries.json  # 200 hand-verified query -> filename pairs├── data/                  # index.npy, meta.sqlite, thumbs/├── Dockerfile└── requirements.txt

The separation that matters is assistant/ knowing nothing about HTTP and api/ knowing nothing about models. It means you can run the indexer as a batch job, test search without starting a server, and swap Gradio for anything else without touching the core.

Step 1 — Environment

Bash
python -m venv .venv && source .venv/bin/activatepip install torch torchvision transformers accelerate pillow \            numpy fastapi "uvicorn[standard]" gradio pytestpython -c "import torch; print(torch.cuda.is_available(), torch.__version__)"

If that prints False, everything still works — indexing 52,000 images on CPU takes roughly 35 minutes instead of 90 seconds, and the generative model will be slow enough that you should use a 2B checkpoint or skip that feature while developing.

Step 2 — The core

Python
# assistant/encoder.pyimport torch, torch.nn.functional as Ffrom transformers import CLIPModel, CLIPProcessorclass Encoder:    def __init__(self, model_id="openai/clip-vit-base-patch32", device=None):        self.model_id = model_id        self.device = device or ("cuda" if torch.cuda.is_available() else "cpu")        self.model = CLIPModel.from_pretrained(model_id).to(self.device).eval()        self.proc = CLIPProcessor.from_pretrained(model_id)    @torch.no_grad()    def encode_images(self, pil_images, batch_size=64):        out = []        for i in range(0, len(pil_images), batch_size):            px = self.proc(images=pil_images[i:i + batch_size],                           return_tensors="pt").to(self.device)            # transformers 5+: the projected embedding is .pooler_output            v = self.model.get_image_features(**px).pooler_output            out.append(F.normalize(v, dim=-1).half().cpu())        return torch.cat(out)    @torch.no_grad()    def encode_texts(self, texts):        tok = self.proc(text=texts, return_tensors="pt",                        padding=True, truncation=True).to(self.device)        v = self.model.get_text_features(**tok).pooler_output        return F.normalize(v, dim=-1).half().cpu()

Every embedding is L2-normalised on the way out, so the dot product is cosine similarity and search is one matrix multiply. Skip that normalisation and long vectors win regardless of direction; your top result becomes whichever image happened to get a large-magnitude embedding, and the bug is invisible because the results are still plausible.

Python
# assistant/index.pyimport json, numpy as np, torchclass Index:    def __init__(self, vectors: np.ndarray, ids: list[str], model_id: str):        assert vectors.shape[0] == len(ids)        self.vectors, self.ids, self.model_id = vectors, ids, model_id    def save(self, path_prefix: str):        np.save(path_prefix + ".npy", self.vectors)        with open(path_prefix + ".json", "w") as f:            json.dump({"ids": self.ids, "model_id": self.model_id,                       "dim": int(self.vectors.shape[1])}, f)    @classmethod    def load(cls, path_prefix: str, expected_model_id: str):        vecs = np.load(path_prefix + ".npy")        meta = json.load(open(path_prefix + ".json"))        if meta["model_id"] != expected_model_id:            raise RuntimeError(                f"Index built with {meta['model_id']} but running "                f"{expected_model_id}. Re-index before searching.")        return cls(vecs, meta["ids"], meta["model_id"])    def search(self, query_vec: np.ndarray, k=10, min_score=0.0):        scores = self.vectors @ query_vec.astype(self.vectors.dtype).ravel()        k = min(k, scores.shape[0])        top = np.argpartition(-scores, k - 1)[:k]        top = top[np.argsort(-scores[top])]        return [(self.ids[i], float(scores[i])) for i in top                if float(scores[i]) >= min_score]

Two details earn their place. The model_id check turns the silent-mismatch disaster into a loud startup error. And np.argpartition finds the top k in O(n) rather than sorting all 52,000 scores — small here, but it is free and it is the right habit.

Zero-shot tagging with prompt ensembling

Python
# assistant/ingest.py (excerpt)LABELS = {    "indoor":       ["a photo taken indoors", "an indoor scene"],    "outdoor":      ["a photo taken outdoors", "an outdoor scene"],    "group":        ["a group photo of many people", "a crowd of people"],    "portrait":     ["a close-up portrait of one person"],    "black_white":  ["a black and white photograph", "a monochrome photo"],    "dancing":      ["people dancing", "a dance floor"],    "speech":       ["a person giving a speech", "someone making a toast"],}def build_label_vectors(encoder):    """Average several prompts per label, then re-normalise."""    import torch.nn.functional as F    names, vecs = [], []    for name, prompts in LABELS.items():        v = encoder.encode_texts(prompts).float()   # (n_prompts, dim)        vecs.append(F.normalize(v.mean(0), dim=-1))        names.append(name)    return names, torch.stack(vecs).numpy().astype("float16")

The double normalisation is not redundant. Averaging unit vectors yields something shorter than unit length, so without the second normalize the similarity scale drifts per label — and labels whose prompts disagree with each other get systematically penalised, which shows up as a category that mysteriously never fires.

Prompt ensembling is worth the two extra lines. The CLIP authors measured roughly 5 percentage points of ImageNet zero-shot accuracy from prompt engineering and ensembling combined, which is larger than the gap between several model sizes.

The thresholding trap

You will want to write if score > 0.3: tag it. Do not hardcode that. Raw cosine similarities from a contrastive model are not calibrated — a score of 0.28 may be excellent for one prompt set and mediocre for another, and it shifts with the model checkpoint. Instead, hand-label 100 images per tag, sweep the threshold, and pick the point that hits your target precision:

Python
def choose_threshold(scores, labels, target_precision=0.90):    """scores: model scores. labels: 1 if the tag truly applies."""    order = np.argsort(-scores)    tp = 0    best = 1.0    for rank, i in enumerate(order, start=1):        tp += labels[i]        if tp / rank >= target_precision:            best = float(scores[i])          # lowest score still above target    return best

Step 3 — The API

Python
# api/server.pyfrom fastapi import FastAPI, HTTPExceptionfrom fastapi.responses import FileResponsefrom pydantic import BaseModelfrom assistant.encoder import Encoderfrom assistant.index import Indexfrom assistant import config, vqaapp = FastAPI()encoder = Encoder(config.CLIP_MODEL_ID)index = Index.load(config.INDEX_PREFIX, config.CLIP_MODEL_ID)class SearchRequest(BaseModel):    query: str    k: int = 10    min_score: float = 0.20@app.post("/search")def search(req: SearchRequest):    if not req.query.strip():        raise HTTPException(400, "empty query")    qv = encoder.encode_texts([req.query]).numpy()    hits = index.search(qv, k=min(req.k, 50), min_score=req.min_score)    return {"results": [{"id": i, "score": round(s, 4)} for i, s in hits]}class AskRequest(BaseModel):    image_id: str    question: str@app.post("/ask")def ask(req: AskRequest):    path = config.path_for(req.image_id)    if path is None:        raise HTTPException(404, "unknown image")    return {"answer": vqa.ask(path, req.question)}@app.get("/image/{image_id}")def image(image_id: str):    path = config.thumb_for(image_id)    if path is None:        raise HTTPException(404, "unknown image")    return FileResponse(path, media_type="image/jpeg")

Note that /search and /ask are separate endpoints with wildly different latency profiles — 50 ms versus 1,200 ms. Merging them, so that every search also generated an answer, would make the fast path as slow as the slow path. Keep them apart and let the UI show a spinner only where one is warranted.

min_score defaults to a real value rather than zero because top-k always returns k results. Search for "a photograph of the surface of Mars" in a wedding archive and, with no floor, you will get ten confident-looking wedding photos. A floor lets the honest answer — nothing — come back.

Step 4 — The interface

Python
# ui/app.pyimport gradio as gr, requestsAPI = "http://localhost:8000"def do_search(query, k):    r = requests.post(f"{API}/search", json={"query": query, "k": int(k)})    hits = r.json()["results"]    return [(f"{API}/image/{h['id']}", f"{h['score']:.3f}") for h in hits]with gr.Blocks() as demo:    gr.Markdown("### Photo assistant")    with gr.Row():        q = gr.Textbox(label="Describe the photo you want", scale=4)        k = gr.Slider(4, 40, value=12, step=4, label="results")    gallery = gr.Gallery(columns=4, height=560)    q.submit(do_search, [q, k], gallery)    with gr.Row():        img_id = gr.Textbox(label="Image id")        question = gr.Textbox(label="Ask about this photo", scale=3)    answer = gr.Textbox(label="Answer")    question.submit(        lambda i, t: requests.post(f"{API}/ask",            json={"image_id": i, "question": t}).json()["answer"],        [img_id, question], answer)demo.launch()

Step 5 — Testing that means something

Unit tests on the index are necessary but easy. The test that actually tells you whether the system works is a retrieval regression suite: a fixed file of hand-verified query-to-image pairs, scored on every change.

Python
# tests/test_index.pyimport json, numpy as npdef test_recall_at_k(index, encoder, k=5):    cases = json.load(open("tests/eval_queries.json"))    hits = 0    ranks = []    for case in cases:        qv = encoder.encode_texts([case["query"]]).numpy()        got = [i for i, _ in index.search(qv, k=50)]        if case["image_id"] in got[:k]:            hits += 1        ranks.append(got.index(case["image_id"]) + 1                     if case["image_id"] in got else 51)    recall = hits / len(cases)    mrr = sum(1.0 / r for r in ranks) / len(ranks)    print(f"Recall@{k} = {recall:.3f}   MRR = {mrr:.3f}")    assert recall >= 0.75          # ratchet this up; never let it fall

Work the numbers so you know what you are looking at. With 200 test queries and the correct image inside the top 5 for 166 of them, Recall@5 = 166/200 = 0.83. If four sample queries return the target at ranks 1, 3, 2 and 10, then MRR = (1 + 0.333 + 0.5 + 0.1)/4 = 0.483. Track both: Recall@10 can look healthy while the right answer sits at position 8, where nobody ever scrolls.

Without a regression suite you have no way to tell an improvement from a regression, and a vector search system degrades silently — nothing errors, results merely become slightly worse, and nobody reports it for months.

Failure modes you will actually hit

symptomcausefix
Every query returns the same handful of photosEmbeddings not normalised, so magnitude dominates directionL2-normalise on write; assert abs(norm − 1) < 1e-3 in the indexer
Results are plausible but never rightIndex built with one checkpoint, queries encoded with anotherThe model_id check in Index.load
Nonsense queries return confident resultsTop-k always returns kA calibrated min_score floor
Indexing crashes at image 31,000A corrupt file, a CMYK JPEG, or an EXIF-rotated imageImage.open(p).convert("RGB") inside a try/except; log and skip
Photos come back sidewaysEXIF orientation ignored, so CLIP saw a rotated imagePIL.ImageOps.exif_transpose before encoding and before thumbnailing
Every .CR2 file is skippedPillow cannot decode camera RAW formatsRead the embedded JPEG preview or decode with rawpy, then encode and thumbnail that
Searching for "SKU-4471-B" failsSemantic embeddings cannot do exact tokensAdd a keyword index and fuse rankings; this is what hybrid search is for
GPU out of memory after a few hundred imagesMissing torch.no_grad(), so an autograd graph is retained per imageDecorate every inference function
Search latency creeps to secondsRe-loading the model or the index per requestLoad once at module import, as in the server above

Deployment

Bash
# DockerfileFROM python:3.11-slimWORKDIR /appRUN apt-get update && apt-get install -y --no-install-recommends \    libgl1 libglib2.0-0 && rm -rf /var/lib/apt/lists/*COPY requirements.txt .RUN pip install --no-cache-dir -r requirements.txtCOPY . .ENV HF_HOME=/modelsEXPOSE 8000CMD ["uvicorn", "api.server:app", "--host", "0.0.0.0", "--port", "8000"]

Two decisions here save real pain. Mount /models as a volume so the container does not re-download several gigabytes of weights on every restart — a cold start that downloads a 7B checkpoint takes minutes and will time out your orchestrator's health check. And keep data/ on a mounted volume too, so rebuilding the image does not throw away a 35-minute index.

Definition of done

areathe bar
Search qualityRecall@5 at or above 0.75 on a 200-query hand-verified suite
Latencyp95 search under 300 ms; question answering has its own endpoint and its own expectations
RobustnessIndexing survives corrupt files, CMYK JPEGs and EXIF rotation without stopping
HonestyAn off-corpus query returns zero results, not ten confident wrong ones
ReproducibilityModel ids pinned; index refuses to load against a different checkpoint
SeparationCore logic importable and testable without starting a web server

Once that holds, the interesting extensions are all cheap. Find similar photos is the same search with an image vector as the query instead of a text vector — a three-line change, since both live in the same space. Faceted filtering combines the stored zero-shot tag scores with the similarity ranking. Duplicate detection is a threshold on pairwise cosine similarity, which on a wedding archive of burst shots will find thousands. Multilingual search needs only a text encoder aligned to the same image space; the image index never changes.

What the project is really teaching

The models here are downloads. The engineering is everything else, and it generalises well beyond photographs.

You will have learned that the split between ingest-time and query-time work is the first architectural decision and it flows from arithmetic, not taste — a 1,200 ms caption in a 300 ms budget is not a tuning problem, it is a scheduling one. You will have learned that embeddings are only meaningful relative to the model that produced them, which is why the version check is load-bearing rather than defensive. You will have learned that unnormalised vectors and uncalibrated thresholds fail quietly, producing plausible results that are wrong, which is the most expensive kind of bug because nothing alerts you.

And you will have a regression suite of 200 verified query-image pairs. That artefact outlives the code. Every future change — a new checkpoint, a re-index, a prompt tweak, an approximate index swapped in when the archive reaches two million photos — gets measured against it. Systems built on vector similarity do not break loudly. They drift. The suite is the only thing that notices.