Multimodal Vision-Language Models

Course Content

Multimodal Vision-Language Models

3 sections · 5 lessons

Visual Question Answering


A supermarket chain runs a shelf-audit pipeline. A field rep photographs a shelf, and the system answers one question: "How many boxes of the own-brand cereal are on the second shelf from the top?"

The team wires up a good open vision-language model, tries it on 200 photos, and gets 61% exact-match accuracy. Not usable. So they dig into the failures, and what they find is that the errors are not one problem — they are four different problems wearing the same costume.

  • On 14 photos the model answered a different question: it counted all the cereal boxes, ignoring "second shelf from the top".
  • On 22 photos it found the right shelf but miscounted — said 6 where there were 8, because two boxes were half-occluded.
  • On 19 photos it could not tell own-brand from a competitor, because that distinction lives in 8-point text on the packaging and the model sees the whole shelf downscaled to a few hundred pixels.
  • On 23 photos the answer was arguably right but written as "about half a dozen", which the exact-match scorer counted as wrong.

Four failures, four completely different fixes. That is the thing to understand about visual question answering before you build one: it is not a task, it is a stack of tasks, and your accuracy number is the product of how well each layer works.

Four problems hiding inside one shelf-audit questionHow many boxeson shelf two?Parse: what isactually being askedLocate: whichshelf is the secondDetect: find every box on itRecognise: isit the own brandCount: a number,not a description
The model can fail any one of the four and still answer fluently, which is why grounded answers beat a bare number.

The four sub-problems hiding inside one question

Every VQA query, however casual, forces the system through four stages. If any stage fails the answer is wrong, and the failure looks identical from the outside.

stagewhat it doesfailure looks like
1. Question understandingParse what is being asked and what shape the answer takes — a number, a colour, a yes/no, a nameA fluent, confident answer to a question nobody asked
2. GroundingLocate the region of the image the question is aboutRight kind of answer, drawn from the wrong part of the picture
3. PerceptionRecognise, read, count or measure what is in that regionOff-by-two counts; misread text; colour named from the background
4. Reasoning & expressionCombine visual evidence with world knowledge, and phrase it in the expected formCorrect understanding stated in a form your scorer rejects

A worked example

Photo: a wooden kitchen table. On it, a glass of orange juice, a plate with two slices of toast, and a folded newspaper. Question: "Is the drink on the left of the plate alcoholic?"

  1. Understanding. This is a yes/no question. The subject is "the drink", constrained by "on the left of the plate". The property asked about is "alcoholic" — which is not visible, it is inferred.
  2. Grounding. Find the plate. Then find a drink whose horizontal position is smaller. If there were two glasses this constraint would be doing real work; models routinely ignore it.
  3. Perception. The glass contains an opaque orange liquid with pulp, in a straight-sided tumbler, next to breakfast food, in daylight.
  4. Reasoning. Orange juice is not alcoholic. Note that nothing in the image says this. The model must supply the fact that orange juice contains no alcohol. Answer: "no".

Stage 4 is why VQA needs a language model rather than a classifier. The answer is not in the pixels. It is in the pixels plus everything the model knows about the world.

Visual question answering is not "reading the answer off the image". It is grounding a question in an image and then reasoning with knowledge the image does not contain.

Not all questions are equally hard

Before you promise anyone an accuracy number, know which category your questions fall into. The spread is enormous.

question typeexampledifficultywhy
Existence / binary"Is there a dog?"EasyGlobal evidence; a coarse image encoding suffices
Attribute"What colour is the car?"EasySingle region, single property
Object identification"What breed is that?"ModerateNeeds fine-grained recognition; degrades on rare classes
Spatial relations"Is the cup left of the laptop?"HardPatch embeddings encode position weakly; left/right is near chance on cluttered scenes
Counting"How many chairs?"HardNo counting mechanism exists; reliability collapses past about 4–5 items
Text in image (OCR)"What is the total on the receipt?"HardResolution-bound — at 336px each patch covers ~14 pixels, smaller than the glyphs
Comparison"Which slice is bigger?"HardTwo groundings plus a relative judgement
External knowledge"What year was this building finished?"Very hardRequires recognising a specific entity and recalling a fact; hallucination-prone
Causal / counterfactual"What happens if she lets go?"Very hardPhysical simulation from a still frame

The supermarket team's question was counting plus spatial constraint plus OCR — three of the hardest categories stacked. 61% was not a model failure. It was a task-design failure.

Anatomy of a working VQA system

A production VQA system is a model surrounded by six pieces of machinery that matter as much as the model does.

Text
image ──► validate ──► preprocess ──► prompt build ──► VLM ──► parse ──► confidence ──► route          (format,      (resize,       (template,               (extract   (score)      (answer /           size,         crop,          answer-format            typed                   review /           EXIF)         enhance)       instruction)             value)                  refuse)

Skip the parse step and you are string-matching free-form prose. Skip the confidence step and you have no way to tell a certain answer from a guess, which means every downstream consumer must treat all answers as unreliable.

The baseline implementation

Python
import osimport torchfrom PIL import Imagefrom transformers import AutoProcessor, AutoModelForImageTextToText# Any chat VLM on the Hugging Face Hub works here; pin the one you evaluated.MODEL_ID = os.environ.get("VLM_MODEL", "Qwen/Qwen3-VL-8B-Instruct")processor = AutoProcessor.from_pretrained(MODEL_ID)model = AutoModelForImageTextToText.from_pretrained(    MODEL_ID, dtype=torch.bfloat16, device_map="auto").eval()ANSWER_STYLE = {    "short":  "Answer with a single word or number, nothing else.",    "yesno":  "Answer only 'yes' or 'no'.",    "count":  "Answer with a single integer. If you cannot count them "              "reliably, answer 'unsure'.",    "long":   "Answer in one or two sentences.",}@torch.inference_mode()def ask(image: Image.Image, question: str, style: str = "short",        max_new_tokens: int = 32) -> str:    messages = [{        "role": "user",        "content": [            {"type": "image"},            {"type": "text", "text": f"{question}\n{ANSWER_STYLE[style]}"},        ],    }]    prompt = processor.apply_chat_template(messages, add_generation_prompt=True)    inputs = processor(images=image, text=prompt,                       return_tensors="pt").to(model.device)    out = model.generate(**inputs, max_new_tokens=max_new_tokens,                         do_sample=False)          # greedy: facts, not prose    new = out[0][inputs["input_ids"].shape[-1]:]    return processor.decode(new, skip_special_tokens=True).strip()

The model id comes from configuration because open VLMs are replaced every few months; AutoModelForImageTextToText loads LLaVA-NeXT, Qwen-VL, Gemma and most other chat VLMs through the same code. If you call a hosted multimodal model instead — the OpenAI, Anthropic Claude and Google Gemini APIs all accept images in a chat message — everything below about prompts, answer formats and parsing still applies, though some hosted reasoning models ignore or reject sampling settings such as temperature. Three deliberate choices there. do_sample=False, because a factual question has one answer and sampling only adds ways to be wrong. A tight max_new_tokens, because a VQA answer is short and the default of 512 makes every rambling answer cost you 512 tokens of latency. And an explicit answer-format instruction — without it, "How many chairs?" returns "There appear to be four chairs visible in the image, although one is partially obscured", and your integer parser returns nothing.

Notice the count style permits "unsure". Giving the model a licence to decline is the single cheapest accuracy improvement available, because the alternative is that it invents a number.

Three techniques that actually move the numbers

1. Question decomposition

Complex questions fail because all four stages must succeed simultaneously. Break the question into steps and each step is easy.

Instead of "How many boxes of own-brand cereal are on the second shelf from the top?", ask in sequence:

Python
steps = [    "How many horizontal shelves are visible? Answer with an integer.",    "Describe only the second shelf from the top. List the product types on it.",    "On that shelf, which products carry the store's own-brand label?",    "How many units of those own-brand products are on that shelf? "    "Answer with an integer, or 'unsure'.",]context = ""for step in steps:    answer = ask(image, context + step, style="long", max_new_tokens=96)    context += f"Q: {step}\nA: {answer}\n"print(context)

This trades latency for accuracy: four forward passes instead of one. It also gives you an inspectable trace — when the final answer is wrong you can see which step went wrong, which is impossible with a single call. The honest caveat is that errors compound. If each of four steps is 90% reliable and they are independent, end-to-end reliability is 0.9⁴ = 65.6%, which can be worse than a single well-prompted call. Decompose when the sub-steps are genuinely much easier than the whole, not by reflex.

2. Grounded VQA — make it show its work

An answer with no evidence cannot be checked. Grounding means asking the model to name the region it used, so a human or a downstream check can verify it.

Python
GROUNDED = """Answer the question, then state exactly what in the imagesupports your answer. If the supporting evidence is not clearly visible,say so instead of guessing.Reply as JSON: {"answer": ..., "evidence": ..., "clearly_visible": true/false}Question: %s"""raw = ask(image, GROUNDED % question, style="long", max_new_tokens=160)

The clearly_visible flag is doing more work than the evidence string. It converts "the model always answers" into "the model can report that the evidence is absent", and that flag is what you route on.

The stronger version pairs the VLM with an object detector: run detection first, feed the detected boxes and labels into the prompt as text, and ask the VLM to answer using those boxes. Counting and spatial questions improve substantially because the detector supplies exactly the two things a VLM is weakest at — discrete object instances and their coordinates.

3. Self-consistency sampling for confidence

Greedy decoding gives one answer and no sense of how fragile it was. Sample the same question several times at moderate temperature and count the votes.

Python
from collections import Counter@torch.inference_mode()def ask_with_confidence(image, question, style="short", k=5, temperature=0.7):    messages = [{"role": "user", "content": [        {"type": "image"},        {"type": "text", "text": f"{question}\n{ANSWER_STYLE[style]}"}]}]    prompt = processor.apply_chat_template(messages, add_generation_prompt=True)    inputs = processor(images=image, text=prompt,                       return_tensors="pt").to(model.device)    out = model.generate(**inputs, max_new_tokens=32, do_sample=True,                         temperature=temperature, top_p=0.9,                         num_return_sequences=k)    cut = inputs["input_ids"].shape[-1]    answers = [processor.decode(o[cut:], skip_special_tokens=True).strip().lower()               for o in out]    top, votes = Counter(answers).most_common(1)[0]    return top, votes / k, answers

Suppose five samples for a counting question return ["3", "3", "4", "3", "2"]. The majority answer is "3" with agreement 3/5 = 0.6. Compare a colour question returning ["red"] × 5 — agreement 1.0. Same model, same call, but the first answer deserves review and the second does not.

The accuracy gain is real and computable. If a single greedy pass is correct 60% of the time and the errors are roughly independent, a 5-way majority vote is correct whenever at least 3 samples are right:

P=(53)(0.6)3(0.4)2+(54)(0.6)4(0.4)+(0.6)5P = \binom{5}{3}(0.6)^3(0.4)^2 + \binom{5}{4}(0.6)^4(0.4) + (0.6)^5

= 10(0.216)(0.16) + 5(0.1296)(0.4) + 0.0778 = 0.3456 + 0.2592 + 0.0778 = 0.683.

That is roughly +8 points for 5× the compute. The independence assumption is optimistic — a model that misreads the packaging will misread it every time, so correlated errors eat into the gain. Self-consistency helps most where the model is genuinely wavering and least where it is confidently wrong. Its real value is the confidence signal, not the accuracy bump.

A VQA system without a confidence score is a system that cannot say "I don't know", and a system that cannot say "I don't know" will confidently make things up in exactly the cases that matter most.

Measuring it properly

Exact string match is the wrong scorer, and it is the reason the supermarket team's 23 "about half a dozen" answers were graded as failures.

The VQA accuracy metric

The standard benchmark metric handles genuine human disagreement. Each question carries answers from 10 annotators, and a predicted answer scores:

Acc(a)=min⁡ ⁣(number of annotators who gave a3, 1)\text{Acc}(a) = \min\!\left(\frac{\text{number of annotators who gave } a}{3},\ 1\right)

averaged over the 10 leave-one-out subsets of 9 annotators. Work through a real case. Question: "What colour is the sea?" Annotators: green ×5, teal ×3, blue ×2. Model predicts "blue".

  • For the 2 subsets where a "blue" annotator is left out, the remaining count is 1: min(1/3, 1) = 0.333
  • For the 8 subsets where a non-"blue" annotator is left out, the count is 2: min(2/3, 1) = 0.667

Score = (2 × 0.333 + 8 × 0.667) / 10 = (0.667 + 5.333) / 10 = 0.60. Partial credit for a defensible minority answer — which is exactly right, because the sea really is arguably blue.

ANLS for text-heavy images

For receipts, forms and documents, near-misses matter. ANLS uses normalised edit distance: score = 1 − (Levenshtein distance / length of the longer string), and anything below 0.5 scores zero.

Gold answer 47.99, model says 4799. One deletion, longer string is 5 characters: 1 − 1/5 = 0.8. Exact match would give 0. Gold 47.99, model says 17.99: one substitution, 1 − 1/5 = 0.8 — and here the leniency is dangerous, because a wrong leading digit on an invoice is not a near-miss. Pick the metric to match the consequence.

metricuse forweakness
Exact matchYes/no, closed label sets, integersPunishes correct paraphrases; needs aggressive normalisation
VQA accuracy (10 annotators)Open-ended natural questionsExpensive to annotate; still surface-level string matching
ANLSDocument and receipt VQAForgives digit errors that change the meaning entirely
LLM-as-judgeLong descriptive answersCostly, needs its own validation against human labels
Accuracy at fixed coverageAnything with a human review pathNeeds a calibrated confidence score to exist

That last row is the one that matters commercially. Reporting "72% accurate" is far less useful than "at a 0.8 agreement threshold we auto-answer 55% of questions at 94% accuracy, and route the remaining 45% to a human". The second statement is something a business can staff against.

What this means for the shelf audit

Go back to the supermarket. Nothing about the model needed to change. What needed to change was the shape of the task, and each fix maps to one of the four stages.

failure (of 200)stagefix
14 — counted the whole shelf unitUnderstandingDecompose: locate the shelf first, then ask about it in a second call
22 — miscounted occluded boxesPerceptionDetector supplies boxes; VLM classifies each crop; count the crops
19 — could not read the brandPerceptionCrop each detected box and re-ask at native resolution instead of on the downscaled shelf
23 — "about half a dozen"ExpressionAnswer-format instruction plus a typed parser; add an "unsure" option

Three of the four fixes are prompt and pipeline engineering, not modelling. This is the general shape of VQA work.

The same architecture, retargeted, is what powers the two big deployed uses. In accessibility, a blind user photographs a medicine packet and asks "what is the dose?" — here a wrong answer is far worse than "I can't read that clearly, try a closer photo", so the confidence threshold is set high and refusal is a first-class outcome. In content moderation, VQA supplements a classifier by answering context questions a classifier cannot — "is this person in medical distress?", "does the overlaid text target a protected group?" — and the design constraint flips: recall matters more than precision, thresholds run low, and everything uncertain goes to a reviewer.

Build it with three habits and it will hold up. Always give the model an explicit way to decline, because a model with no exit will fabricate. Always produce a confidence number alongside the answer, because without one you cannot set a routing threshold and every consumer must assume the worst. And always test with the grey-image check — run your questions against a blank grey rectangle, and if the answers barely change, your visual pathway is broken and every accuracy number you have measured so far is measuring your language prior.