Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Vision: documents, screenshots and scanned forms
Three requests arrived in the same month, and all three were about things PolicyPal could not see. First, the retrieval misses from Section 3: 23 of the 400 policies were scanned PDFs, old signed documents like the 2019 relocation policy, with no text layer at all. pypdf returned empty strings, so those policies had never been searchable. Second, about 11% of IT questions in the pilot arrived with a screenshot: "What does this error mean?" with a picture of a dialog box, and no text. Third, HR operations asked for help with the medical certificates employees submit for sick leave longer than two days: about 400 a month, many handwritten, each typed in by hand in about four minutes.
Modern models can read images directly. A vision-capable model accepts an image, a photo, a screenshot or a scanned page, as part of the input, alongside text, and can describe, transcribe or extract from it. That opens all three problems. It also brings new costs, new ways to be wrong, and, with medical certificates, some of the most sensitive data in the company.
This lesson handles each request with a different pattern: transcription for search, reading for understanding, and extraction with review for data entry.
Two ways to read an image
Classic OCR, then a text model
- Runs locally and cheaply, for example with Tesseract
- Excellent on clean, printed, single-column text
- Loses layout: headings, tables and columns blur together
- Weak on handwriting and on screenshots
A vision-capable model
- Reads layout, tables, handwriting and interface screenshots
- Can transcribe, answer or extract in one step
- About 1,500 to 2,500 input tokens per page image
- A hosted model means the image leaves your network
The team tried Tesseract first on the 23 scanned policies, because it was free and local. The words came out mostly right, but section numbers merged into paragraphs and tables turned into streams of numbers, which undid every chunking lesson from Section 3. On 15 test questions about those policies, recall@5 was 0.40. With a vision model transcribing each page into Markdown, it was 0.87. For documents with structure, layout is not a detail; it is the meaning.
Transcribing scanned policies for the index
The approach is to render each page to an image, ask the model to transcribe it faithfully into Markdown, and send the result through the same heading-aware chunker as every other policy. Rendering uses pypdfium2; the image goes to the model as an image content block, which the LLM client passes through unchanged.
1# policypal/vision.py2import base643import io45import pypdfium2 as pdfium67TRANSCRIBE = ("Transcribe this scanned policy page into Markdown. Keep section numbers "8 "with their headings and keep tables as Markdown tables. Do not summarise, "9 "correct or add anything. Write [illegible] for text you cannot read.")1011def page_images(path: str, scale: float = 2.0):12 pdf = pdfium.PdfDocument(path)13 for i in range(len(pdf)):14 image = pdf[i].render(scale=scale).to_pil() # scale 2 is about 144 dpi15 buf = io.BytesIO()16 image.save(buf, format="PNG")17 yield i + 1, base64.standard_b64encode(buf.getvalue()).decode()1819def image_block(b64: str, media_type: str = "image/png") -> dict:20 return {"type": "image", "source": {"type": "base64", "media_type": media_type, "data": b64}}2122def transcribe_pdf(llm, path: str) -> list[tuple[int, str]]:23 pages = []24 for number, b64 in page_images(path):25 content = [image_block(b64), {"type": "text", "text": TRANSCRIBE}]26 reply = llm.complete("You transcribe scanned documents exactly.",27 [{"role": "user", "content": content}], max_tokens=2000)28 pages.append((number, reply.text))29 return pagesOne page per call keeps failures small: a bad page can be retried alone, and page numbers stay exact for citations. The instruction "do not summarise, correct or add anything" matters more than it looks. Without it, the model sometimes tidied a policy's wording or filled a smudged number with a plausible one, and a plausible number in a policy is exactly the failure this whole course is about. [illegible] gives it an honest option instead.
The cost is small because it happens once. The 23 policies had about 310 pages. At about 2,000 input tokens per page image and 700 output tokens of Markdown, that is 620,000 input and 217,000 output tokens, or about $3.40 at mid-tier prices. A person then spot-checked 10% of the pages against the scans and found two numbers to fix. Chunks from transcribed pages carry source_type: "scan", and their citations say "transcribed from a scanned document", so employees know to click through for anything that matters.
Reading screenshots in IT questions
For a screenshot, PolicyPal sends the image with the question in the same user message. The model reads the error text, for example "VPN connection failed. Error 809", and the agent then searches the policies for it, where BM25 finds the troubleshooting section by the exact code. A screenshot costs about 1,600 input tokens, roughly $0.003, and answer correctness on 50 screenshot questions was 84%, against 41% when the user had to retype the error.
Screenshots bring two risks. They often show more than the error: chat windows, email subjects, other people's names. PolicyPal tells users to crop, does not store images longer than seven days, and never adds image content to the answer cache. And a screenshot can contain text written to manipulate the model, such as a fake dialog saying "ignore your rules". Text in an image is data, exactly like text in a question, and Section 9's defences apply to it too.
Extracting fields from scanned forms
Medical certificates are a different kind of task. The goal is not an answer but data: a handful of fields to pre-fill the HR system, which a person then approves. That calls for the structured-output pattern from Section 2, with one addition: an honest way for the model to say a field is unreadable.
1# policypal/forms.py2from datetime import date3from difflib import SequenceMatcher45from pydantic import BaseModel, ConfigDict67class MedicalCertificate(BaseModel):8 model_config = ConfigDict(extra="forbid")9 patient_name: str10 doctor_name: str11 doctor_registration: str # "" when unreadable12 issued_on: str # "YYYY-MM-DD", "" when unreadable13 rest_from: str14 rest_to: str15 unreadable_fields: list[str] # names of fields the model could not read1617def review_reasons(cert: MedicalCertificate, employee_name: str) -> list[str]:18 reasons = [f"unreadable: {f}" for f in cert.unreadable_fields]19 try:20 start, end = date.fromisoformat(cert.rest_from), date.fromisoformat(cert.rest_to)21 if end < start:22 reasons.append("rest ends before it starts")23 if (end - start).days > 30:24 reasons.append("rest longer than 30 days")25 except ValueError:26 reasons.append("rest dates not readable")27 if SequenceMatcher(None, cert.patient_name.lower(), employee_name.lower()).ratio() < 0.8:28 reasons.append("name does not match the employee")29 return reasonsThe schema has no diagnosis field, on purpose. HR needs to know that a registered doctor certified rest between two dates. It does not need the medical condition, so PolicyPal never extracts it, and the prompt says not to mention it. Collecting only what the process needs is the simplest privacy control there is.
review_reasons applies the checks code can do well: dates that parse and make sense, a plausible length and a name that matches the employee who uploaded it. Any reason at all puts the certificate in a review queue with the reasons shown. With no reasons, the fields pre-fill the form, and an HR person still approves it. Nothing is approved automatically.
On a test set of 400 past certificates, fields were exactly right 93% of the time on printed certificates and 78% on handwritten ones. With the review rules, 29% of certificates were flagged for review. Of the 71% that were not, a spot check found a wrong field in 1.2%, and HR's approval step caught those. Average handling time fell from about four minutes to about one.
Certificates are health data, so the team made the processing decision explicitly, with the security and legal teams. Extraction runs on the hosted model under the company's data-processing agreement: no training on the data and limited retention at the provider. Images are deleted from PolicyPal's storage once HR approves or rejects. Had that agreement not been possible, the fallback was a self-hosted open-weight vision model, which the team had tested at about 6 points lower field accuracy.
Check your understanding
0 of 3 answered
1.Tesseract read most words on the scanned policies correctly, yet recall@5 on questions about them was only 0.40. Why?
2.Why does the medical certificate schema have no field for the diagnosis?
3.A certificate's extracted rest_to date is before its rest_from date. What should PolicyPal do?