Course Content
Capstone Project: Multimodal Assistant
1 sections · 6 lessons
Text & Image Processing
You want the assistant to look at a photo of your fridge and suggest dinner. The obvious first attempt is to read the file and hand the bytes to the chat model:
1with open("fridge.jpg", "rb") as f:2 raw = f.read()34llm.invoke([{"role": "user", "content": raw}])5# TypeError: Object of type bytes is not JSON serializableFine, encode it. Now you get a 200 response and a reply that says: "I'd be happy to help! However, I don't see an image — could you describe what's in your fridge?" You just paid for 5,300 tokens of base64 gibberish and received an apology. Pasted into the text of a message, an image is only a long string of characters, and the model reads it as text. Vision-capable models do exist — most current models from the major providers accept images — but only through a dedicated image part of the request, and a model without an image encoder cannot be prompted into seeing.
The second tempting mistake is subtler. You get a caption from somewhere, paste it into a giant f-string, and ship it. Two weeks later there are five near-identical prompt strings scattered across three modules — one for captions, one for questions about images, one for plain chat, two for variants nobody remembers writing. You fix a typo in the instruction "answer concisley" and it stays broken in four places.
This stage fixes both. You build a vision pipeline that turns pixels into something a language model can actually consume, and a prompt layer that gives every prompt in the project exactly one home.
The pipeline that makes multimodal work
Here is the mechanism you are building: a caption pipeline. A native vision model turns the image into tokens inside the model itself; in this design, by contrast, the language model never sees pixels. Something else converts the image into text, and from that point onward it is an ordinary text problem. It runs locally and free, and it makes every step visible.
fridge.jpg | | PIL: load, convert to RGB, resize to 224x224 v +-----------------------------+ | VisionHandler | | | | CLIP -----> 512-dim vector, comparable with text vectors | "does this image match the label 'a fridge'?" | | | BLIP -----> "a refrigerator with milk, eggs and a bunch | of carrots on the middle shelf" +-----------------------------+ | | caption (plain text) v +-----------------------------+ | PromptBuilder (what to say) + LLMChainManager (how to call) +-----------------------------+ | v "You could make a carrot and egg fried rice..."Two different models, doing two different jobs. CLIP compares — it places an image and a piece of text in the same vector space so you can ask "how well does this image match this description?". BLIP describes — it generates a sentence from pixels. You need both, for different reasons, and confusing them is the most common conceptual error here.
The language model in a multimodal assistant is not doing the seeing. Its input is a sentence someone else wrote about the image, and every limitation of that sentence becomes a limitation of the answer.
Prompts are data, not code
Two responsibilities get conflated constantly: what to say to the model and how to call the model. Split them and both become testable. PromptBuilder is pure string manipulation with zero network dependencies — you can unit test every prompt with no real API key and no internet. LLMChainManager owns the client, the temperature, the retries, and the error handling.
1# src/text_processor.py2from langchain_openai import ChatOpenAI3from config.settings import settings4from src.exceptions import TextProcessingError5from src.utils import logger678class PromptBuilder:9 """Every prompt the assistant sends lives here. No exceptions."""1011 SYSTEM_BASE = (12 "You are a careful, knowledgeable assistant. "13 "You state your reasoning before your conclusion. "14 "If you are unsure, you say so rather than guessing."15 )1617 @staticmethod18 def system_prompt(context: str = "") -> str:19 if not context:20 return PromptBuilder.SYSTEM_BASE21 return f"{PromptBuilder.SYSTEM_BASE}\n\nRelevant context:\n{context}"2223 @staticmethod24 def reasoning_prompt(query: str) -> str:25 # The numbered steps are load-bearing, not decoration: models are26 # measurably more reliable on multi-step tasks when the prompt27 # supplies the decomposition instead of asking for a bare answer.28 return (29 f"Answer this question step by step.\n\n"30 f"Question: {query}\n\n"31 "1. State what is being asked.\n"32 "2. Break the problem into parts.\n"33 "3. Work through each part.\n"34 "4. State your final answer on its own line, prefixed 'Answer:'.\n"35 )3637 @staticmethod38 def image_analysis_prompt(caption: str, question: str) -> str:39 # The model receives a DESCRIPTION, never pixels. Saying so in the40 # prompt stops it inventing detail the caption never mentioned.41 return (42 "An automatic captioning model produced this description of "43 f"an image:\n\n\"{caption}\"\n\n"44 f"Question about the image: {question}\n\n"45 "Answer using only what the description supports. If the "46 "description does not contain enough information, say exactly "47 "what is missing rather than speculating."48 )That last instruction is the difference between a useful assistant and a confident liar. Given the caption "a refrigerator with milk, eggs and carrots" and the question "is the milk expired?", an unconstrained model will happily invent a date. Constrained, it says the description does not mention a date. Prompt discipline is how you keep a captioning model's blind spots visible instead of laundering them into fluent prose.
The call layer
1class LLMChainManager:2 """The only object in the project that holds an LLM client."""34 def __init__(self):5 self.llm = ChatOpenAI(6 model=settings.text_model,7 temperature=settings.temperature,8 max_tokens=settings.max_tokens,9 api_key=settings.openai_api_key,10 timeout=30,11 max_retries=2,12 )13 logger.info("LLM ready: %s (temperature=%s)",14 settings.text_model, settings.temperature)1516 def _invoke(self, prompt: str, system: str | None = None) -> str:17 messages = []18 if system:19 messages.append({"role": "system", "content": system})20 messages.append({"role": "user", "content": prompt})21 try:22 response = self.llm.invoke(messages)23 except Exception as e:24 logger.error("LLM call failed: %s", e)25 raise TextProcessingError(f"LLM call failed: {e}") from e26 logger.debug("LLM returned %d chars", len(response.content))27 return response.content.strip()2829 def chat(self, message: str, context: str = "") -> str:30 return self._invoke(message, PromptBuilder.system_prompt(context))3132 def reason(self, query: str) -> str:33 return self._invoke(PromptBuilder.reasoning_prompt(query))3435 def analyse_image(self, caption: str, question: str) -> str:36 return self._invoke(37 PromptBuilder.image_analysis_prompt(caption, question)38 )Every public method routes through _invoke, which means retries, timeouts, logging, and error wrapping are written once. Add a token counter or a cost meter later and you add it in one place. Scatter llm.invoke() calls across six modules instead and you will add it in six, and forget the seventh.
People often ask what a "chain" buys over calling the API directly. Honestly, for a single call: very little. The value shows up at the second and third call, when you want the same retry policy, the same model choice, and the same failure logging everywhere. A chain is a named, single-responsibility unit of LLM work. That is worth having even if the implementation is thirty lines of your own code rather than a framework's.
CLIP: images and text in one vector space
CLIP was trained on hundreds of millions of image-caption pairs with one objective: make a matching image and caption produce vectors that point in the same direction, and non-matching pairs point apart. The result is an embedding space where you can compare a picture to a sentence with plain arithmetic.
The comparison is cosine similarity: the cosine of the angle between two vectors, which ignores their lengths and measures only direction.
Work it through in four dimensions so the mechanics are visible. Let the image vector be a=[0.6,0.8,0,0], which has length 0.36+0.64=1. Compare it to two candidate text vectors, both also unit length:
- b=[0,1,0,0]: dot product =0.6×0+0.8×1=0.8, so similarity 0.8.
- c=[0.8,−0.6,0,0]: dot product =0.6×0.8+0.8×(−0.6)=0.48−0.48=0, so similarity 0 — perfectly unrelated.
1# src/vision_handler.py2import torch3from PIL import Image4from transformers import CLIPModel, CLIPProcessor5from src.exceptions import VisionError6from src.utils import logger789class ImageEmbedder:10 """CLIP: image and text into one comparable 512-dim space."""1112 def __init__(self, model_name: str = "openai/clip-vit-base-patch32"):13 self.device = "cuda" if torch.cuda.is_available() else "cpu"14 logger.info("Loading CLIP on %s", self.device)15 self.model = CLIPModel.from_pretrained(model_name).to(self.device)16 self.processor = CLIPProcessor.from_pretrained(model_name)17 self.model.eval()1819 def load_image(self, path: str) -> Image.Image:20 try:21 img = Image.open(path)22 except (OSError, ValueError) as e:23 raise VisionError(f"cannot open {path}: {e}") from e24 # Convert unconditionally: CMYK, palette, and RGBA images all25 # crash downstream tensor code that assumes 3 channels.26 if img.mode != "RGB":27 logger.warning("Converting %s from %s to RGB", path, img.mode)28 img = img.convert("RGB")29 return img3031 @torch.no_grad()32 def match(self, path: str, labels: list[str]) -> list[tuple[str, float]]:33 """Rank candidate labels by how well they describe the image."""34 img = self.load_image(path)35 inputs = self.processor(36 text=labels, images=img, return_tensors="pt", padding=True37 ).to(self.device)3839 out = self.model(**inputs)40 # logits_per_image is cosine similarity multiplied by CLIP's41 # learned logit_scale, which is close to 100.42 probs = out.logits_per_image.softmax(dim=1)[0]43 ranked = sorted(44 zip(labels, probs.tolist()), key=lambda x: -x[1]45 )46 logger.debug("CLIP ranking for %s: %s", path, ranked)47 return rankedWhy the raw numbers look disappointing
Run CLIP on a genuinely matching image and caption and you get a cosine similarity around 0.30, not 0.95. Beginners see that and conclude the model failed. It did not. CLIP's contrastive training compresses everything into a narrow band; what carries the signal is the gap between candidates, not the absolute value.
Watch what CLIP's logit_scale of roughly 100 does to a small gap. Suppose three labels score cosine similarities of 0.31, 0.24 and 0.19 against your fridge photo. Multiply by 100 to get logits of 31, 24 and 19, then take the softmax. Subtracting the maximum first for numerical stability, the exponentials are e0=1, e−7=0.000912, and e−12=0.0000061. The sum is 1.000918, so the probabilities are:
| Label | Cosine similarity | Logit (× 100) | Softmax probability |
|---|---|---|---|
| "a photo of a refrigerator interior" | 0.31 | 31 | 0.9991 |
| "a photo of a kitchen counter" | 0.24 | 24 | 0.00091 |
| "a photo of a car engine" | 0.19 | 19 | 0.0000061 |
A cosine gap of 0.07 — which looks like noise — becomes a probability ratio of about 1,100 to 1. That scale factor is exactly why you should never threshold on raw cosine values ("accept if above 0.8" rejects everything) and should instead compare candidates against each other.
With CLIP, ranking is trustworthy and absolute scores are not. Design every feature around "which of these labels fits best", never around "does this exceed 0.8".
Resize before you do anything else
CLIP ViT-B/32 and most caption models consume 224×224 pixels. A modern phone photo is 4032×3024, which is 12,192,768 pixels. As a float32 tensor with three channels that is 12,192,768×3×4=146 MB. The 224×224 version is 224×224×3×4=602,112 bytes, about 0.6 MB — roughly 243 times smaller. The processor resizes for you, but if you build your own batching, forgetting this is how a batch of eight images turns into a 1.2 GB allocation and an out-of-memory kill.
Captioning: pixels to a sentence
CLIP can rank labels you supply, but it cannot invent a description. For that you need a captioning model — an image encoder bolted to a small text decoder, trained to generate the caption a human would write.
1from transformers import BlipForConditionalGeneration, BlipProcessor234class ImageCaptioner:5 def __init__(self, model_name: str = "Salesforce/blip-image-captioning-base"):6 self.device = "cuda" if torch.cuda.is_available() else "cpu"7 self.processor = BlipProcessor.from_pretrained(model_name)8 self.model = BlipForConditionalGeneration.from_pretrained(9 model_name10 ).to(self.device)11 self.model.eval()1213 @torch.no_grad()14 def caption(self, img: Image.Image, prompt: str | None = None) -> str:15 inputs = self.processor(img, text=prompt, return_tensors="pt")16 inputs = inputs.to(self.device)17 ids = self.model.generate(**inputs, max_new_tokens=40, num_beams=3)18 text = self.processor.decode(ids[0], skip_special_tokens=True).strip()19 if not text:20 raise VisionError("captioner returned an empty string")21 logger.info("Caption: %s", text)22 return text232425class VisionHandler:26 """The single entry point for anything image-shaped."""2728 def __init__(self):29 self.embedder = ImageEmbedder()30 self.captioner = ImageCaptioner()3132 def describe(self, path: str) -> str:33 return self.captioner.caption(self.embedder.load_image(path))3435 def classify(self, path: str, labels: list[str]) -> tuple[str, float]:36 ranked = self.embedder.match(path, labels)37 return ranked[0]Now understand what you have and have not got. A captioner produces a generic description of the salient content. It is not visual question answering. Given a photo of a receipt and the question "what was the total?", a base captioning model returns something like "a piece of paper on a table". The information you wanted was never extracted, so no amount of LLM reasoning downstream can recover it.
| Captioning (BLIP) | CLIP matching | A native vision LLM | |
|---|---|---|---|
| Input | Image | Image + candidate labels | Image + arbitrary question |
| Output | One generic sentence | Ranked scores over your labels | Free-form answer |
| Answers "what is the total on this receipt?" | No | No | Often yes |
| Answers "is this indoors or outdoors?" | Sometimes | Yes, if you supply both labels | Yes |
| Runs offline | Yes | Yes | Usually no |
| Cost per image | CPU time only | CPU time only | Per-token API charge |
Build the caption-plus-LLM pipeline because it teaches the mechanism and runs free on a laptop. Know its ceiling, and know that swapping VisionHandler.describe() for a call to a native vision model is a one-method change precisely because the boundary is clean.
Joining the two halves
1class MultimodalResponder:2 def __init__(self, vision: VisionHandler, text: LLMChainManager):3 self.vision = vision4 self.text = text56 def answer_about_image(self, path: str, question: str) -> dict:7 caption = self.vision.describe(path)8 answer = self.text.analyse_image(caption, question)9 # Return the caption too: when the answer is wrong, you need to10 # know whether the caption was wrong or the reasoning was.11 return {"caption": caption, "answer": answer}Returning the intermediate caption alongside the answer is not decoration. It is the single most useful debugging affordance in the whole vision path, because it splits "the model got it wrong" into two distinct, separately fixable failures.
The tool placeholder, and why eval is not acceptable
To exercise the pipeline end to end you need the model to be able to do something beyond talk — a calculator, say. The quick version is to have the model emit a JSON blob in a fenced block, split the string, and evaluate the expression:
def calculate(expression: str) -> str: return str(eval(expression)) # DO NOT SHIP THISConsider what happens when the model, prompted by a user who typed something adversarial, emits __import__('os').system('rm -rf ~') as the expression. eval executes it. You have handed arbitrary code execution to a text generator that can be steered by anyone who can type into your assistant. The minimum acceptable version restricts what can run:
1import ast2import operator34_OPS = {5 ast.Add: operator.add, ast.Sub: operator.sub,6 ast.Mult: operator.mul, ast.Div: operator.truediv,7 ast.Pow: operator.pow, ast.USub: operator.neg,8}91011def safe_calculate(expression: str) -> float:12 """Evaluate arithmetic only. No names, no calls, no attributes."""1314 def walk(node):15 if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):16 return node.value17 if isinstance(node, ast.BinOp) and type(node.op) in _OPS:18 left, right = walk(node.left), walk(node.right)19 if isinstance(node.op, ast.Pow) and abs(right) > 100:20 raise ValueError("exponent too large") # 9**9**9 would hang21 return _OPS[type(node.op)](left, right)22 if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:23 return _OPS[type(node.op)](walk(node.operand))24 raise ValueError(f"disallowed expression element: {type(node).__name__}")2526 # round away binary float noise: 84.50 * 0.15 is 12.67499999999999927 return round(walk(ast.parse(expression, mode="eval").body), 10)This parses the expression into a syntax tree and walks it, permitting only numeric literals and six arithmetic operators. Anything else — a function call, a name lookup, an attribute access — raises before it can run. The exponent cap closes the one hole left: 9**9**9 is pure arithmetic, and computing it would hang the process. The final round hides binary floating-point noise, so the tip example later returns 12.675 rather than 12.674999999999999. The string-splitting on fenced JSON that feeds it is still fragile, and a schema-validated tool registry supersedes it once the reasoning engine exists; the sandboxing, though, is not optional at any stage.
Tests worth writing here
1# tests/test_text_processor.py2from src.text_processor import PromptBuilder345def test_prompt_includes_the_question():6 p = PromptBuilder.reasoning_prompt("What is 12 percent of 350?")7 assert "12 percent of 350" in p8 assert "Answer:" in p91011def test_image_prompt_forbids_speculation():12 p = PromptBuilder.image_analysis_prompt("a red bicycle", "what colour?")13 assert "a red bicycle" in p14 assert "speculat" in p.lower()151617def test_system_prompt_omits_empty_context():18 assert "Relevant context" not in PromptBuilder.system_prompt("")19 assert "Relevant context" in PromptBuilder.system_prompt("user likes tea")Every one of those runs in milliseconds with no real API key (a dummy value in .env satisfies the settings check), because PromptBuilder makes no network calls. That is the payoff of the split. For the vision side, test the deterministic parts — that a CMYK image is converted to RGB, that a missing file raises VisionError and not FileNotFoundError, that match() returns labels sorted descending. Do not assert on the exact text of a caption; generation is not deterministic across versions and such a test will fail for reasons that have nothing to do with your code.
When things go wrong here
| Symptom | Cause | Fix |
|---|---|---|
| Model replies "I don't see an image" | Bytes or base64 pasted into the message text | Caption first and send the caption text; or, with a vision-capable model, send the image as an image part of the request |
RuntimeError: Input type (torch.FloatTensor) and weight type (torch.cuda.FloatTensor) should be the same | Model moved to GPU, inputs left on CPU | inputs.to(self.device) on every processor output, not just the model |
| Caption is empty or a single word | Grayscale, CMYK or RGBA input silently mangled | img.convert("RGB") before the processor, always |
| All CLIP scores between 0.20 and 0.32, "nothing matches" | Expecting absolute similarity near 1.0 | Compare candidates by rank; use softmax over logits, never a fixed cosine threshold |
| Process killed during batch captioning | Full-resolution tensors — 146 MB per 12 MP image | Resize to 224×224 before batching; cap batch size |
| First call takes 40 s, later calls are fast | Model weights downloading and loading | Instantiate VisionHandler once at startup and reuse it; never construct it per request |
| Answer confidently states facts the caption never mentioned | Prompt permits speculation | The "use only what the description supports" clause, plus returning the caption for inspection |
| Every prompt tweak requires editing three files | Prompt strings inlined in logic | Move them into PromptBuilder; logic modules should contain no prompt text at all |
Acceptance criteria for this stage
grep -rn '"""' src/*.py | grep -ci "you are a"finds prompt text only insidetext_processor.py— zero prompt strings anywhere else.VisionHandler().classify("examples/cat.jpg", ["a photo of a cat", "a photo of a car"])ranks the cat label first with a softmax probability above 0.9.- The same call on a deliberately CMYK-converted copy of that image succeeds and logs one WARNING about colour conversion.
describe()on three different sample images returns three different non-empty captions, each under 40 tokens.answer_about_image()returns a dict containing bothcaptionandanswer, and asking about something absent from the caption produces an explicit "the description does not say" rather than an invented fact.safe_calculate("2 + 3 * 4")returns 14;safe_calculate("__import__('os')")raisesValueError.- The prompt tests pass with a dummy
OPENAI_API_KEYand the network switched off — proof the prompt layer has no hidden network dependency. - Second and subsequent image requests complete in under 2 seconds, confirming models are loaded once rather than per call.
What this buys you for the rest of the build
The concrete win is that every later capability plugs into a text pipe. Speech becomes text before it reaches the reasoning layer. Retrieved memories arrive as text in the context slot that system_prompt() already accepts. A tool result is a string appended to a prompt. Because you refused to let pixels leak past VisionHandler and refused to let prompt strings leak out of PromptBuilder, adding a modality later means writing one adapter, not rewriting the assistant.
The second win is diagnostic. When a user complains the assistant got their photo wrong, you have two artefacts, not one: the caption and the answer. If the caption says "a piece of paper on a table", the vision model is your problem and a better captioner or a native vision endpoint is the fix. If the caption is accurate and the answer still misses, the prompt is your problem and it is a one-method edit. Systems without that split produce a single opaque failure and an afternoon of guessing — which is precisely the cost the boundary was drawn to avoid.