Course Content
Multimodal Vision-Language Models
3 sections · 5 lessons
Visual Question Answering
A supermarket chain runs a shelf-audit pipeline. A field rep photographs a shelf, and the system answers one question: "How many boxes of the own-brand cereal are on the second shelf from the top?"
The team wires up a good open vision-language model, tries it on 200 photos, and gets 61% exact-match accuracy. Not usable. So they dig into the failures, and what they find is that the errors are not one problem — they are four different problems wearing the same costume.
- On 14 photos the model answered a different question: it counted all the cereal boxes, ignoring "second shelf from the top".
- On 22 photos it found the right shelf but miscounted — said 6 where there were 8, because two boxes were half-occluded.
- On 19 photos it could not tell own-brand from a competitor, because that distinction lives in 8-point text on the packaging and the model sees the whole shelf downscaled to a few hundred pixels.
- On 23 photos the answer was arguably right but written as "about half a dozen", which the exact-match scorer counted as wrong.
Four failures, four completely different fixes. That is the thing to understand about visual question answering before you build one: it is not a task, it is a stack of tasks, and your accuracy number is the product of how well each layer works.
The four sub-problems hiding inside one question
Every VQA query, however casual, forces the system through four stages. If any stage fails the answer is wrong, and the failure looks identical from the outside.
| stage | what it does | failure looks like |
|---|---|---|
| 1. Question understanding | Parse what is being asked and what shape the answer takes — a number, a colour, a yes/no, a name | A fluent, confident answer to a question nobody asked |
| 2. Grounding | Locate the region of the image the question is about | Right kind of answer, drawn from the wrong part of the picture |
| 3. Perception | Recognise, read, count or measure what is in that region | Off-by-two counts; misread text; colour named from the background |
| 4. Reasoning & expression | Combine visual evidence with world knowledge, and phrase it in the expected form | Correct understanding stated in a form your scorer rejects |
A worked example
Photo: a wooden kitchen table. On it, a glass of orange juice, a plate with two slices of toast, and a folded newspaper. Question: "Is the drink on the left of the plate alcoholic?"
- Understanding. This is a yes/no question. The subject is "the drink", constrained by "on the left of the plate". The property asked about is "alcoholic" — which is not visible, it is inferred.
- Grounding. Find the plate. Then find a drink whose horizontal position is smaller. If there were two glasses this constraint would be doing real work; models routinely ignore it.
- Perception. The glass contains an opaque orange liquid with pulp, in a straight-sided tumbler, next to breakfast food, in daylight.
- Reasoning. Orange juice is not alcoholic. Note that nothing in the image says this. The model must supply the fact that orange juice contains no alcohol. Answer: "no".
Stage 4 is why VQA needs a language model rather than a classifier. The answer is not in the pixels. It is in the pixels plus everything the model knows about the world.
Visual question answering is not "reading the answer off the image". It is grounding a question in an image and then reasoning with knowledge the image does not contain.
Not all questions are equally hard
Before you promise anyone an accuracy number, know which category your questions fall into. The spread is enormous.
| question type | example | difficulty | why |
|---|---|---|---|
| Existence / binary | "Is there a dog?" | Easy | Global evidence; a coarse image encoding suffices |
| Attribute | "What colour is the car?" | Easy | Single region, single property |
| Object identification | "What breed is that?" | Moderate | Needs fine-grained recognition; degrades on rare classes |
| Spatial relations | "Is the cup left of the laptop?" | Hard | Patch embeddings encode position weakly; left/right is near chance on cluttered scenes |
| Counting | "How many chairs?" | Hard | No counting mechanism exists; reliability collapses past about 4–5 items |
| Text in image (OCR) | "What is the total on the receipt?" | Hard | Resolution-bound — at 336px each patch covers ~14 pixels, smaller than the glyphs |
| Comparison | "Which slice is bigger?" | Hard | Two groundings plus a relative judgement |
| External knowledge | "What year was this building finished?" | Very hard | Requires recognising a specific entity and recalling a fact; hallucination-prone |
| Causal / counterfactual | "What happens if she lets go?" | Very hard | Physical simulation from a still frame |
The supermarket team's question was counting plus spatial constraint plus OCR — three of the hardest categories stacked. 61% was not a model failure. It was a task-design failure.
Anatomy of a working VQA system
A production VQA system is a model surrounded by six pieces of machinery that matter as much as the model does.
image ──► validate ──► preprocess ──► prompt build ──► VLM ──► parse ──► confidence ──► route (format, (resize, (template, (extract (score) (answer / size, crop, answer-format typed review / EXIF) enhance) instruction) value) refuse)Skip the parse step and you are string-matching free-form prose. Skip the confidence step and you have no way to tell a certain answer from a guess, which means every downstream consumer must treat all answers as unreliable.
The baseline implementation
1import os2import torch3from PIL import Image4from transformers import AutoProcessor, AutoModelForImageTextToText56# Any chat VLM on the Hugging Face Hub works here; pin the one you evaluated.7MODEL_ID = os.environ.get("VLM_MODEL", "Qwen/Qwen3-VL-8B-Instruct")8processor = AutoProcessor.from_pretrained(MODEL_ID)9model = AutoModelForImageTextToText.from_pretrained(10 MODEL_ID, dtype=torch.bfloat16, device_map="auto"11).eval()1213ANSWER_STYLE = {14 "short": "Answer with a single word or number, nothing else.",15 "yesno": "Answer only 'yes' or 'no'.",16 "count": "Answer with a single integer. If you cannot count them "17 "reliably, answer 'unsure'.",18 "long": "Answer in one or two sentences.",19}2021@torch.inference_mode()22def ask(image: Image.Image, question: str, style: str = "short",23 max_new_tokens: int = 32) -> str:24 messages = [{25 "role": "user",26 "content": [27 {"type": "image"},28 {"type": "text", "text": f"{question}\n{ANSWER_STYLE[style]}"},29 ],30 }]31 prompt = processor.apply_chat_template(messages, add_generation_prompt=True)32 inputs = processor(images=image, text=prompt,33 return_tensors="pt").to(model.device)34 out = model.generate(**inputs, max_new_tokens=max_new_tokens,35 do_sample=False) # greedy: facts, not prose36 new = out[0][inputs["input_ids"].shape[-1]:]37 return processor.decode(new, skip_special_tokens=True).strip()The model id comes from configuration because open VLMs are replaced every few months; AutoModelForImageTextToText loads LLaVA-NeXT, Qwen-VL, Gemma and most other chat VLMs through the same code. If you call a hosted multimodal model instead — the OpenAI, Anthropic Claude and Google Gemini APIs all accept images in a chat message — everything below about prompts, answer formats and parsing still applies, though some hosted reasoning models ignore or reject sampling settings such as temperature. Three deliberate choices there. do_sample=False, because a factual question has one answer and sampling only adds ways to be wrong. A tight max_new_tokens, because a VQA answer is short and the default of 512 makes every rambling answer cost you 512 tokens of latency. And an explicit answer-format instruction — without it, "How many chairs?" returns "There appear to be four chairs visible in the image, although one is partially obscured", and your integer parser returns nothing.
Notice the count style permits "unsure". Giving the model a licence to decline is the single cheapest accuracy improvement available, because the alternative is that it invents a number.
Three techniques that actually move the numbers
1. Question decomposition
Complex questions fail because all four stages must succeed simultaneously. Break the question into steps and each step is easy.
Instead of "How many boxes of own-brand cereal are on the second shelf from the top?", ask in sequence:
1steps = [2 "How many horizontal shelves are visible? Answer with an integer.",3 "Describe only the second shelf from the top. List the product types on it.",4 "On that shelf, which products carry the store's own-brand label?",5 "How many units of those own-brand products are on that shelf? "6 "Answer with an integer, or 'unsure'.",7]89context = ""10for step in steps:11 answer = ask(image, context + step, style="long", max_new_tokens=96)12 context += f"Q: {step}\nA: {answer}\n"13print(context)This trades latency for accuracy: four forward passes instead of one. It also gives you an inspectable trace — when the final answer is wrong you can see which step went wrong, which is impossible with a single call. The honest caveat is that errors compound. If each of four steps is 90% reliable and they are independent, end-to-end reliability is 0.9⁴ = 65.6%, which can be worse than a single well-prompted call. Decompose when the sub-steps are genuinely much easier than the whole, not by reflex.
2. Grounded VQA — make it show its work
An answer with no evidence cannot be checked. Grounding means asking the model to name the region it used, so a human or a downstream check can verify it.
1GROUNDED = """Answer the question, then state exactly what in the image2supports your answer. If the supporting evidence is not clearly visible,3say so instead of guessing.45Reply as JSON: {"answer": ..., "evidence": ..., "clearly_visible": true/false}67Question: %s"""89raw = ask(image, GROUNDED % question, style="long", max_new_tokens=160)The clearly_visible flag is doing more work than the evidence string. It converts "the model always answers" into "the model can report that the evidence is absent", and that flag is what you route on.
The stronger version pairs the VLM with an object detector: run detection first, feed the detected boxes and labels into the prompt as text, and ask the VLM to answer using those boxes. Counting and spatial questions improve substantially because the detector supplies exactly the two things a VLM is weakest at — discrete object instances and their coordinates.
3. Self-consistency sampling for confidence
Greedy decoding gives one answer and no sense of how fragile it was. Sample the same question several times at moderate temperature and count the votes.
1from collections import Counter23@torch.inference_mode()4def ask_with_confidence(image, question, style="short", k=5, temperature=0.7):5 messages = [{"role": "user", "content": [6 {"type": "image"},7 {"type": "text", "text": f"{question}\n{ANSWER_STYLE[style]}"}]}]8 prompt = processor.apply_chat_template(messages, add_generation_prompt=True)9 inputs = processor(images=image, text=prompt,10 return_tensors="pt").to(model.device)11 out = model.generate(**inputs, max_new_tokens=32, do_sample=True,12 temperature=temperature, top_p=0.9,13 num_return_sequences=k)14 cut = inputs["input_ids"].shape[-1]15 answers = [processor.decode(o[cut:], skip_special_tokens=True).strip().lower()16 for o in out]17 top, votes = Counter(answers).most_common(1)[0]18 return top, votes / k, answersSuppose five samples for a counting question return ["3", "3", "4", "3", "2"]. The majority answer is "3" with agreement 3/5 = 0.6. Compare a colour question returning ["red"] × 5 — agreement 1.0. Same model, same call, but the first answer deserves review and the second does not.
The accuracy gain is real and computable. If a single greedy pass is correct 60% of the time and the errors are roughly independent, a 5-way majority vote is correct whenever at least 3 samples are right:
= 10(0.216)(0.16) + 5(0.1296)(0.4) + 0.0778 = 0.3456 + 0.2592 + 0.0778 = 0.683.
That is roughly +8 points for 5× the compute. The independence assumption is optimistic — a model that misreads the packaging will misread it every time, so correlated errors eat into the gain. Self-consistency helps most where the model is genuinely wavering and least where it is confidently wrong. Its real value is the confidence signal, not the accuracy bump.
A VQA system without a confidence score is a system that cannot say "I don't know", and a system that cannot say "I don't know" will confidently make things up in exactly the cases that matter most.
Measuring it properly
Exact string match is the wrong scorer, and it is the reason the supermarket team's 23 "about half a dozen" answers were graded as failures.
The VQA accuracy metric
The standard benchmark metric handles genuine human disagreement. Each question carries answers from 10 annotators, and a predicted answer scores:
averaged over the 10 leave-one-out subsets of 9 annotators. Work through a real case. Question: "What colour is the sea?" Annotators: green ×5, teal ×3, blue ×2. Model predicts "blue".
- For the 2 subsets where a "blue" annotator is left out, the remaining count is 1: min(1/3, 1) = 0.333
- For the 8 subsets where a non-"blue" annotator is left out, the count is 2: min(2/3, 1) = 0.667
Score = (2 × 0.333 + 8 × 0.667) / 10 = (0.667 + 5.333) / 10 = 0.60. Partial credit for a defensible minority answer — which is exactly right, because the sea really is arguably blue.
ANLS for text-heavy images
For receipts, forms and documents, near-misses matter. ANLS uses normalised edit distance: score = 1 − (Levenshtein distance / length of the longer string), and anything below 0.5 scores zero.
Gold answer 47.99, model says 4799. One deletion, longer string is 5 characters: 1 − 1/5 = 0.8. Exact match would give 0. Gold 47.99, model says 17.99: one substitution, 1 − 1/5 = 0.8 — and here the leniency is dangerous, because a wrong leading digit on an invoice is not a near-miss. Pick the metric to match the consequence.
| metric | use for | weakness |
|---|---|---|
| Exact match | Yes/no, closed label sets, integers | Punishes correct paraphrases; needs aggressive normalisation |
| VQA accuracy (10 annotators) | Open-ended natural questions | Expensive to annotate; still surface-level string matching |
| ANLS | Document and receipt VQA | Forgives digit errors that change the meaning entirely |
| LLM-as-judge | Long descriptive answers | Costly, needs its own validation against human labels |
| Accuracy at fixed coverage | Anything with a human review path | Needs a calibrated confidence score to exist |
That last row is the one that matters commercially. Reporting "72% accurate" is far less useful than "at a 0.8 agreement threshold we auto-answer 55% of questions at 94% accuracy, and route the remaining 45% to a human". The second statement is something a business can staff against.
What this means for the shelf audit
Go back to the supermarket. Nothing about the model needed to change. What needed to change was the shape of the task, and each fix maps to one of the four stages.
| failure (of 200) | stage | fix |
|---|---|---|
| 14 — counted the whole shelf unit | Understanding | Decompose: locate the shelf first, then ask about it in a second call |
| 22 — miscounted occluded boxes | Perception | Detector supplies boxes; VLM classifies each crop; count the crops |
| 19 — could not read the brand | Perception | Crop each detected box and re-ask at native resolution instead of on the downscaled shelf |
| 23 — "about half a dozen" | Expression | Answer-format instruction plus a typed parser; add an "unsure" option |
Three of the four fixes are prompt and pipeline engineering, not modelling. This is the general shape of VQA work.
The same architecture, retargeted, is what powers the two big deployed uses. In accessibility, a blind user photographs a medicine packet and asks "what is the dose?" — here a wrong answer is far worse than "I can't read that clearly, try a closer photo", so the confidence threshold is set high and refusal is a first-class outcome. In content moderation, VQA supplements a classifier by answering context questions a classifier cannot — "is this person in medical distress?", "does the overlaid text target a protected group?" — and the design constraint flips: recall matters more than precision, thresholds run low, and everything uncertain goes to a reviewer.
Build it with three habits and it will hold up. Always give the model an explicit way to decline, because a model with no exit will fabricate. Always produce a confidence number alongside the answer, because without one you cannot set a routing threshold and every consumer must assume the worst. And always test with the grey-image check — run your questions against a blank grey rectangle, and if the answers barely change, your visual pathway is broken and every accuracy number you have measured so far is measuring your language prior.