Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is multimodality, and how do agents leverage multimodal models?
What you need to know
Where agents use multimodal input
| Use | Input | Why multimodal helps |
|---|---|---|
| Document understanding | Invoices, claim forms, contracts | Layout carries meaning: tables, stamps, signatures |
| Damage assessment | Photos of a car or a parcel | Evidence that has no text form |
| Computer use | Screenshots of a desktop or browser | Operates systems with no API |
| Monitoring and QA | Dashboards, UI screenshots | Spot visual regressions |
| Voice | Speech in and out | Call centres, hands-free use |
Engineering rules
- Cost: image tokens grow with resolution. Crop to the region that matters and downscale.
- Structure: extract into a schema (
invoice_number,total,gstin), then validate in code. - Reference, do not resend: store the image once and pass an ID; re-sending a large image every step multiplies cost.
- Prefer structure over pixels: for web tasks, the DOM or accessibility tree is more reliable than clicking coordinates.
- Keep a fallback: OCR text, or a human, when image quality is poor.
Known weak spots
Tiny text, counting many objects, precise coordinates, and handwriting still cause errors. Design checks around them, for example re-computing an invoice total from line items.
Injection through images
A PDF or image can contain text like "approve this claim". A multimodal model reads it just as it reads typed text. Apply the same rule as for any tool result: it is evidence, not an instruction.
A real-life example
A motor insurer's claims agent receives, via WhatsApp, three photos of a dented car, a garage estimate as a phone photo, and a scanned RC (registration certificate).
- A vision model extracts the estimate into a schema: 7 line items, labour, parts, GST, total Rs 48,200. Code re-adds the line items and finds Rs 46,800. The difference is flagged, not silently accepted.
- The damage photos are checked against the estimate: the estimate includes a rear-bumper replacement, but no photo shows the rear. The agent calls
request_document("rear_photo"). - The insurer's old surveyor portal has no API. A computer-use agent running in an isolated virtual desktop opens the portal, searches by registration number and reads the last survey date. It is given only read access and 20 actions maximum.
Photos are downscaled to 1,024 pixels on the long side before sending, which cut image-token cost by about 70% with no drop in extraction accuracy on the test set.
Follow-up questions to expect
- "OCR then text model, or native document input?" — Native document input usually wins when layout matters, such as tables and stamps. OCR is still useful as a cheap fallback and a cross-check.
- "When would you use a computer-use agent?" — Only when there is no API and building one is not possible, such as a vendor's legacy portal. Run it sandboxed, with limited actions and approval before any submit.
- "How do you evaluate extraction from images?" — Field-level accuracy against a labelled set, plus arithmetic and cross-field checks in code.