Multimodal Vision-Language Models

Course Content

Multimodal Vision-Language Models

3 sections · 5 lessons

LLaVA and Flamingo Architectures


A radiologist sends you a chest X-ray and asks a simple question: "Is there anything in the lower left lobe, and if so, describe it."

You have a contrastive image-text model — the kind that embeds pictures and sentences into one shared vector space and scores how well they match. So you do the only thing it can do. You write out a list of candidate answers: "a chest X-ray with a nodule in the lower left lobe", "a normal chest X-ray", "a chest X-ray with pleural effusion". The model scores each one and picks a winner.

It works, sort of. But notice what just happened: you wrote the answer. The model only chose from a menu you supplied. It cannot say "there is a 12mm opacity, roughly circular, near the costophrenic angle" because you did not think to type that sentence. Ask it something you did not anticipate and it has no move at all.

A matching model can only ever rank options. To get an answer — free-form text the model composes itself — you need a language model doing the generating, and you need it to be able to see. That is the problem LLaVA and Flamingo solve, in two quite different ways.

Two ways to get an image into a language modelLLaVA: into the prompt• Projection maps patches to token space• 576 image tokens sit before the text• LLM weights unchanged in shape• Every image eats context and costFlamingo: into the layers• Perceiver resamples to 64 latents• Gated cross-attention inside each block• Gate starts at zero, so day one is safe• Interleaves many images in one sequence
Pasting images into the prompt is simpler and quadratically expensive; injecting them into the layers keeps the text sequence short.

What generation actually requires

Contrastive models like CLIP give you two encoders and a similarity score. Everything they can do reduces to "rank these candidates". Powerful, cheap, and fundamentally closed-ended.

capabilitycontrastive matchergenerative VLM
Classify into known labelsYes, and very fastYes, but slower and needs parsing
Retrieve images by textYes — one matmul over a precomputed indexNo; nothing to index
Answer an unanticipated questionNoYes
Describe something in a novel sentenceNoYes
Multi-step reasoning over an imageNoYes, with the LLM's reasoning
Follow instructions ("be brief", "answer in JSON")NoYes

The engineering question is: how do you get pixels into a language model that only accepts token embeddings? Two answers dominate.

Every generative vision-language model is a language model plus a decision about where the visual information enters: through the input sequence, or through the layers.

LLaVA: put the image in the prompt

LLaVA (Large Language and Vision Assistant, 2023) takes the simpler route, and it is simple to the point of feeling like a trick. A language model consumes a sequence of vectors. Image patches, once encoded, are also vectors. So: convert the patches into vectors of the right shape and paste them into the sequence, in front of the user's question.

Text
image ──► CLIP ViT-L/14-336 (FROZEN)              │              │  576 patch vectors, 1024 dims each              ▼         MLP projector (TRAINED)  1024 → 4096 → 4096              │              │  576 vectors, 4096 dims — same shape as a token embedding              ▼   [ v1 v2 ... v576 ][ "USER:" "What" "is" "wrong" "?" ]  ──► Vicuna / Llama LLM                                                              │                                                              ▼                                                     "There is a 12mm opacity..."

The three components, and which one is doing the work

1. The vision encoder. A pretrained CLIP ViT-L/14 at 336×336 pixels. Dividing 336 by the 14-pixel patch size gives 24 across, so 24×24 = 576 patch tokens, each a 1024-dimensional vector. LLaVA uses the patch tokens rather than the pooled summary vector, because a single summary vector cannot support questions about specific regions. The encoder stays frozen throughout.

2. The projector. The vision encoder emits 1024-dim vectors; the LLM's embedding space is 4096-dim for a 7B Llama-class model. Something must bridge them. LLaVA-1.0 used a single linear layer; LLaVA-1.5 upgraded to a two-layer MLP with a GELU in the middle, which measurably improved results.

Count the parameters: 1024×4096 = 4,194,304 for the first layer, plus 4096×4096 = 16,777,216 for the second. About 21 million parameters — roughly 0.3% of a 7-billion-parameter model. That tiny module is the entire bridge between vision and language.

3. The language model. Vicuna, Llama, Mistral, Qwen — an ordinary decoder-only LLM, unmodified. It has no idea that some of its input tokens came from an image. They are just vectors, and it attends to them exactly as it attends to word embeddings.

LLaVA's real claim is that a language model does not need to be taught to see. It needs the visual features handed to it in a dialect it already speaks.

The cost of pasting images into the prompt

This design has a price, and it is worth computing because it drives every practical decision downstream.

576 visual tokens consume 576 slots of context. In a 4,096-token window that is 14% gone before the user types anything. Worse, LLaVA-1.6 introduced "AnyRes" tiling to handle higher resolutions: split the image into four crops plus one downscaled global view, encode each, and concatenate. That is 5 × 576 = 2,880 tokens — 70% of a 4,096-token context for one image. Attention is quadratic in sequence length, so the attention cost of the visual portion rises by (2880/576)² = 25×.

The KV cache is the number that surprises people. For a 7B Llama-class model — 32 layers, hidden size 4096, fp16 — each token stores a key and a value per layer: 2 × 4096 × 2 bytes = 16 KB per layer, × 32 layers = 512 KB per token. So:

configurationvisual tokensKV cache for the image aloneat batch 8
LLaVA-1.5, 336px576288 MB2.3 GB
LLaVA-1.6 AnyRes, 4 tiles + global2,8801.41 GB11.3 GB

On a 24 GB card holding 14 GB of fp16 weights, the second row does not fit at batch 8. This is why people are surprised that a "7B model" will not serve eight concurrent image requests on a consumer GPU. The image, not the weights, is what fills your memory.

Two-stage training, and why the order matters

You cannot just train the whole thing end to end on instruction data. The projector starts random, so its outputs are noise; gradients flowing into a strong pretrained LLM from noise will damage it. LLaVA splits training in two.

Stage 1 — feature alignmentStage 2 — instruction tuning
Data~558K image/caption pairs (short alt-text style)~665K multi-turn visual instruction conversations
Vision encoderFrozenFrozen
ProjectorTrained (~21M params)Trained
LLMFrozenTrained (full fine-tune or LoRA)
GoalMake projector output land where the LLM expects word embeddingsTeach the model to answer questions, follow format instructions, refuse

Stage 1 is cheap — 21 million trainable parameters, one epoch. Stage 2 is where the model learns to behave. The instruction data was itself generated by prompting a text-only LLM with image captions and bounding boxes, then asking it to write questions and answers — a neat bootstrap that made the whole approach affordable. LLaVA-1.5-13B trains in about a day on eight A100s, which is why dozens of variants exist.

The failure mode here is skipping stage 1. Train everything at once from a random projector and you get a model whose language quality has visibly degraded — repetition, broken grammar, lost instruction-following — because the early noisy gradients scrambled the LLM before the visual features meant anything.

Using it

Bash
pip install torch transformers accelerate pillow
Python
import torchfrom PIL import Imagefrom transformers import LlavaNextProcessor, LlavaNextForConditionalGenerationmodel_id = "llava-hf/llava-v1.6-mistral-7b-hf"processor = LlavaNextProcessor.from_pretrained(model_id)model = LlavaNextForConditionalGeneration.from_pretrained(    model_id, dtype=torch.float16, device_map="auto")image = Image.open("xray_0417.png").convert("RGB")# The chat template inserts the special image placeholder in the# right position for this checkpoint. Do not hand-write it.messages = [{    "role": "user",    "content": [        {"type": "image"},        {"type": "text",         "text": "Describe any abnormality in the lower left lobe. "                 "If there is none, say 'none'."},    ],}]prompt = processor.apply_chat_template(messages, add_generation_prompt=True)inputs = processor(images=image, text=prompt, return_tensors="pt").to(model.device)with torch.inference_mode():    out = model.generate(**inputs, max_new_tokens=256, do_sample=False)# Slice off the prompt so you get only the newly generated tokens.generated = out[0][inputs["input_ids"].shape[-1]:]print(processor.decode(generated, skip_special_tokens=True).strip())

Two things in that snippet are the source of most bug reports. First, always use apply_chat_template — every checkpoint has its own image placeholder token and its own turn markers, and hand-writing the prompt string produces silently degraded output rather than an error. Second, slice off the prompt before decoding; generate returns prompt plus completion, and forgetting this makes your model appear to echo the question back.

Flamingo: put the image in the layers

Flamingo (DeepMind, 2022) asks a different question. Suppose you have a very large, very good frozen language model — 70 billion parameters — and you do not want to fine-tune it at all, because doing so is expensive and risks damaging it. And suppose your input is not one image and one question, but a document: text, image, more text, two more images, a caption, a question.

Flamingo does not put visual tokens in the sequence. It leaves the text sequence alone and injects vision between the frozen layers.

The Perceiver Resampler

The first problem is that a video clip or a multi-image document produces a wildly variable number of visual features. Flamingo compresses whatever it gets to a fixed 64 latent vectors.

It works by cross-attention with learned queries. Initialise 64 learnable query vectors. Let them attend to the visual features (which may be 256 patches, or 1,600 patches from five video frames). Each query pulls out whatever it has learned to look for. Output: always exactly 64 vectors, regardless of input size.

Text
variable input                fixed output  256 patches  ─┐ 1600 patches  ─┼──►  cross-attention with 64 learned queries  ──►  64 vectors 4800 patches  ─┘

The saving is real: a 5-frame video at 256 patches per frame is 1,280 features, compressed to 64 — a 20× reduction before anything touches the language model.

Gated cross-attention, and the zero that makes it safe

Between (some of) the frozen LM layers, Flamingo inserts new cross-attention blocks. Text tokens attend to the 64 visual latents. The frozen layers are untouched; the new layers are trained.

The critical detail is the gate. Each inserted block computes:

h←h+tanh⁡(α)⋅CrossAttn(h,visual latents)h \leftarrow h + \tanh(\alpha) \cdot \text{CrossAttn}(h, \text{visual latents})

where α is a learned scalar initialised at zero. Since tanh(0) = 0, at the very first training step the entire visual branch contributes exactly nothing, and the network's output is bit-for-bit identical to the frozen language model's. The gate then opens gradually as training finds visual features useful.

Compare that with the alternative. Insert randomly-initialised cross-attention with no gate, and step one of training pushes random noise into a 70B model that took millions of GPU-hours to produce. Loss spikes, the language model's fluency collapses, and you spend the next week wondering why your VLM writes gibberish.

Initialising a new module so that it starts as the identity function is one of the most reliable tricks in deep learning. Flamingo's zero-initialised tanh gate and LoRA's zero-initialised second matrix are the same idea.

Flamingo-80B pairs a frozen 70B Chinchilla LM with roughly 10 billion new parameters in resamplers and gated cross-attention layers. The LM never moves.

Interleaving and in-context learning

Because vision enters through cross-attention with a per-token mask controlling which image each text token may attend to, Flamingo handles arbitrary interleavings natively:

Text
<image> This is a chinchilla. They are found in the Andes.<image> This is a shiba. They are popular in Japan.<image> This is

The masking rule is that each text token attends only to the most recent preceding image. That gives clean few-shot prompting: show two labelled examples, then a third image, and the model completes the pattern — no gradient updates, no fine-tuning. Flamingo demonstrated that a 32-shot prompt could beat task-specific fine-tuned models on several benchmarks, which was the headline result at the time.

OpenFlamingo and the first Idefics were the open reproductions, trained on interleaved web documents. Tellingly, later Idefics versions switched to the LLaVA-style design. The cross-attention idea lives on in models such as Meta's Llama 3.2 Vision, which adds cross-attention layers to a text model whose original weights stay frozen.

Choosing between the two designs

Contrastive (CLIP-style)LLaVA-style projectorFlamingo-style cross-attention
OutputSimilarity scoreFree textFree text
Where vision entersSeparate towerInput sequenceInserted layers
LLM modified?No LLMYes — fine-tuned in stage 2No — stays frozen
Trainable paramsAll of itProjector + LLMResampler + gated layers only
Context cost per imageNone576–2,880 tokensZero sequence tokens
Multiple / interleaved imagesN/AAwkward; context fills fastNative
Few-shot in-context learningN/AWeakStrong
Implementation effortLowLow — no LLM surgeryHigh — must modify layers
Best forSearch, ranking, filtering, moderationSingle-image Q&A, captioning, OCR-ish tasksVideo, documents, many-image reasoning

In practice the LLaVA pattern won the mainstream. The Qwen-VL family, InternVL, Idefics2/3 and Gemma put visual tokens into the sequence (the newest, such as Qwen3.5, train vision and text together from the start rather than attaching a projector to a finished LLM), usually with a smarter compressor than a plain MLP — merging neighbouring patches, or a small resampler that turns 576 patches into 64 or 256 tokens, which is Flamingo's idea recycled inside LLaVA's architecture. The reason is prosaic: it requires no surgery on the language model, so it rides every improvement to open LLMs for free. Cross-attention survives where sequence length is the binding constraint — long video, multi-page documents.

Sampling, and why VQA is not creative writing

A generative VLM is still a language model, so decoding settings apply — and the right settings for visual question answering are the opposite of the defaults people copy from chat demos.

tasksettingsreason
Factual VQA, extraction, OCRdo_sample=False (greedy)There is one right answer. Sampling only adds ways to be wrong, and makes results non-reproducible.
Structured output (JSON, labels)Greedy, plus constrained decoding if availableAny randomness breaks the parser eventually
Descriptive captioningtemperature=0.6, top_p=0.9Some variation reads better; greedy captions are repetitive and flat
Marketing copy, alt-text variantstemperature=0.9, generate n and pickYou want diversity and a human is choosing

Also cap max_new_tokens aggressively. A VQA answer needs 32 tokens; leaving the default at 512 means that when the model starts rambling — and it will — you pay for 512 tokens of latency to discover it.

The failure mode everyone hits: object hallucination

Ask a generative VLM "what objects are in this kitchen?" and it will confidently list a refrigerator, a microwave and a toaster whether or not the toaster is there. The cause is that the language model's prior — kitchens contain toasters — is strong, and the visual evidence is only 576 vectors competing against everything the LLM learned from text. When the image is ambiguous, the prior wins.

symptomcausefix
Lists objects that are not presentLanguage prior overrides weak visual evidenceAsk yes/no per object; explicitly permit "no" and "not visible" in the prompt
Confidently misreads small text576 tokens at 336px is ~14px per patch; fine print is sub-patchUse a higher-resolution or tiling variant, or crop and re-ask
Gets left/right and counts wrongPatch order carries position weakly; no counting mechanismUse a detector for counts and coordinates; let the VLM describe, not measure
Answers change between runsdo_sample=True left on from a chat exampleGreedy decoding for anything factual
Ignores the image entirelyHand-written prompt string missing the image placeholder tokenUse apply_chat_template; sanity-check with a blank image and see if the answer changes

That last diagnostic is worth building into every pipeline. Run your question against a solid grey image. If the answer is roughly the same as with the real image, the visual pathway is not contributing and you have a plumbing bug, not a model quality problem.

What this means when you deploy one

The X-ray question from the start is now answerable, but the systems reality is different from a matching model in three specific ways.

Memory planning is about tokens, not weights. A 7B model in fp16 is 14 GB, which is the number everyone quotes. But at 2,880 visual tokens the KV cache is 1.41 GB per concurrent request. Your batch size is determined by image resolution, not parameter count. If throughput matters, the highest-leverage change is usually reducing visual tokens — a resampler that compresses 576 patches to 144 cuts KV cache by 4× and attention cost by 16×, usually at a small accuracy loss on non-OCR tasks.

Cache what does not change. If you ask ten questions about the same image, the visual prefill is identical every time. Encoding the image once and reusing the KV cache for the visual prefix removes the dominant cost of every request after the first. On a 576-token image that is roughly 2 × 7e9 × 576 ≈ 8 TFLOPs of prefill saved per repeat question.

Quantisation is nearly free here. 4-bit weight quantisation takes a 7B model from 14 GB to about 4 GB with modest quality loss, and the vision encoder — 300M parameters — can stay in fp16 because it costs nothing. Quantise the LLM, not the encoder.

And the design choice itself follows from the input shape rather than from benchmarks. One image, one question, and you want it working this afternoon: use a LLaVA-style model off the shelf. A twenty-page scanned contract, or thirty seconds of video, or a few labelled examples you want the model to imitate without training: you need compression and cross-attention, and you should reach for the Flamingo lineage. Choosing on leaderboard scores alone is how teams end up trying to fit a video into a 4,096-token window one frame at a time.