Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
An extraction prompt is at 78% field-level accuracy and the business needs 95%, but you cannot fine-tune yet. What do you do?
What you need to know
Going from 78% to 95% means removing about three quarters of the errors. You cannot do that by guessing which sentence of the prompt to change. You need to know what the errors are.
Build the taxonomy
- Freeze a test set — 200 labelled documents that never change during the work.
- Run and collect — every field that does not match the label is one failure.
- Tag each failure — by cause, not by field.
- Count — sort the buckets by size and fix the biggest first.
A typical result for invoices:
| Bucket | Share of failures | Targeted fix |
|---|---|---|
| Format (date as "3rd March", amount with commas) | 30% | Schema-constrained output with typed fields |
| Invented value when field is absent | 25% | Allow null; add an example where the field is missing |
| Two similar fields confused (invoice date vs due date) | 20% | Clear field descriptions; a separate focused call |
| Value from the wrong section (billing vs shipping address) | 15% | Few-shot examples from these failures |
| Genuinely ambiguous document | 10% | Flag for human review |
The fixes
- Constrain the output. Structured-output modes restrict generation to your JSON Schema, so format errors become impossible rather than unlikely. Types such as
dateandDecimalremove a whole bucket. - Few-shot from the errors. Choose 5 to 8 examples that come from the failing buckets, including one where a field is truly absent and the correct output is
null. - Decompose. If two fields are always confused, extract them in their own short call with a focused instruction. Two cheap calls often beat one clever prompt.
- Verify high-value fields. A second check confirms each extracted value appears in the source text. A value that cannot be found is flagged, not trusted.
1class Invoice(BaseModel):2 invoice_number: str3 invoice_date: date = Field(description="Date the invoice was issued, not the due date")4 due_date: date | None = Field(description="Payment due date; null if not printed")5 total_amount: Decimal6 gstin: str | None = None78def verify(inv: Invoice, source: str) -> list[str]:9 return [f for f in ("invoice_number", "gstin")10 if getattr(inv, f) and getattr(inv, f) not in source]The field descriptions are sent to the model, so they carry the disambiguation. verify returns fields whose value does not appear literally in the document.
Measure per field
A 94% average can hide one field at 60%. Report accuracy per field, and set the 95% target per field where the business needs it.
A real-life example
Scenario (illustrative numbers). An accounts-payable team extracts 14 fields from supplier invoices. Average field accuracy is 78%. The engineer labels all 616 failing fields in the 200-document test set.
The biggest bucket is date format (190 failures), and structured output removes almost all of it: 84%. Adding null for absent fields and an example of a missing due date fixes most invented values: 89%. Splitting invoice date and due date into a focused call: 92%. Few-shot examples from the billing-versus-shipping failures, plus a verifier on GSTIN and invoice number: 95.4% average, with every field above 93%. The remaining failures are mostly blurry scans, which now go to a review queue.
Follow-up questions to expect
- "How do you avoid overfitting the prompt to the test set?" — Keep a second held-out set you check only at the end, and refresh it with new documents each month.
- "When would you move to fine-tuning?" — When a bucket survives all prompt fixes. The labelled failures are already the start of a training set.
- "Does a bigger model solve it?" — Sometimes partly, but it costs more on every call; measure the per-bucket gain before paying for it.