Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is multimodal prompting?


What you need to know

Images are cut into small patches, and each patch becomes tokens the model reads like words. Higher resolution means more patches, more tokens, higher cost and — up to a limit set by the provider — better reading of small detail.

Common uses

  • Extracting fields from invoices, receipts, application forms and PDFs
  • Reading charts and dashboards
  • Screenshot to code, or UI bug reports
  • Product photo checks (right item, damage, missing parts)
  • Alt text for accessibility

Prompting practices

  1. Put the image before the question and say what to focus on: "the table in the lower half".
  2. Ask for structured output with a schema.
  3. Give an escape value — "unreadable" — so the model does not invent a blurred digit.
  4. Send enough resolution; crop to the region that matters instead of sending a huge full page.
  5. Several images? Label them ("Image 1: front, Image 2: label") and refer to them by label.

Known limits

  • Small, dense or handwritten text
  • Exact counts of many similar objects
  • Precise positions and measurements
  • Confident reading of text that is actually blurred

What changed

By 2026 frontier models read screenshots, charts and native PDFs well, so a separate OCR step and step-by-step "read the axis labels first" instructions are often unnecessary. Test that on your own documents before removing them, and keep validation — models are better, not perfect.

A real-life example

A B2B marketplace lets small suppliers send invoice photos over WhatsApp. The first prompt with the photo:

Text
[photo of invoice]What's in this invoice?

The reply is a paragraph, and on a blurred photo it states a GSTIN confidently — with two wrong characters. The improved prompt:

Text
[photo of invoice]Extract these fields from the invoice photo above:supplier_name, gstin, invoice_number, invoice_date (YYYY-MM-DD),total_inr (number).If any character of a field is not clearly readable, return"unreadable" for that field. Do not guess characters.

Code then checks that the GSTIN is 15 characters and that its last check character is valid. Blurred photos now return "unreadable" about 70% of the time, and those go to a clerk, while clear photos are processed automatically. Wrong GSTINs in the ledger fall from 3% to 0.2%.

Follow-up questions to expect

  • "How do you reduce the token cost of images?" — Crop to the relevant region, downscale where detail is not needed, and cache repeated images or documents.
  • "PDF: send as images or text?" — Many APIs now accept PDFs directly and use both the text layer and page images; for scanned PDFs the page images are what matter.
  • "How do you evaluate a vision extraction?" — The same way as text extraction: a labelled set of real images, including poor-quality ones, scored per field.