Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is multimodality, and how do agents leverage multimodal models?


What you need to know

Where agents use multimodal input

UseInputWhy multimodal helps
Document understandingInvoices, claim forms, contractsLayout carries meaning: tables, stamps, signatures
Damage assessmentPhotos of a car or a parcelEvidence that has no text form
Computer useScreenshots of a desktop or browserOperates systems with no API
Monitoring and QADashboards, UI screenshotsSpot visual regressions
VoiceSpeech in and outCall centres, hands-free use

Engineering rules

  • Cost: image tokens grow with resolution. Crop to the region that matters and downscale.
  • Structure: extract into a schema (invoice_number, total, gstin), then validate in code.
  • Reference, do not resend: store the image once and pass an ID; re-sending a large image every step multiplies cost.
  • Prefer structure over pixels: for web tasks, the DOM or accessibility tree is more reliable than clicking coordinates.
  • Keep a fallback: OCR text, or a human, when image quality is poor.

Known weak spots

Tiny text, counting many objects, precise coordinates, and handwriting still cause errors. Design checks around them, for example re-computing an invoice total from line items.

Injection through images

A PDF or image can contain text like "approve this claim". A multimodal model reads it just as it reads typed text. Apply the same rule as for any tool result: it is evidence, not an instruction.

A real-life example

A motor insurer's claims agent receives, via WhatsApp, three photos of a dented car, a garage estimate as a phone photo, and a scanned RC (registration certificate).

  • A vision model extracts the estimate into a schema: 7 line items, labour, parts, GST, total Rs 48,200. Code re-adds the line items and finds Rs 46,800. The difference is flagged, not silently accepted.
  • The damage photos are checked against the estimate: the estimate includes a rear-bumper replacement, but no photo shows the rear. The agent calls request_document("rear_photo").
  • The insurer's old surveyor portal has no API. A computer-use agent running in an isolated virtual desktop opens the portal, searches by registration number and reads the last survey date. It is given only read access and 20 actions maximum.

Photos are downscaled to 1,024 pixels on the long side before sending, which cut image-token cost by about 70% with no drop in extraction accuracy on the test set.

Follow-up questions to expect

  • "OCR then text model, or native document input?" — Native document input usually wins when layout matters, such as tables and stamps. OCR is still useful as a cheap fallback and a cross-check.
  • "When would you use a computer-use agent?" — Only when there is no API and building one is not possible, such as a vendor's legacy portal. Run it sandboxed, with limited actions and approval before any submit.
  • "How do you evaluate extraction from images?" — Field-level accuracy against a labelled set, plus arithmetic and cross-field checks in code.