Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

An extraction prompt is at 78% field-level accuracy and the business needs 95%, but you cannot fine-tune yet. What do you do?


Field accuracy after each targeted fix78%84%89%92%95.4%01234schema output: datesnull allowedsplitconfused fieldsfew-shotplus verifier616 failures tagged by cause before any prompt edit.
Each step removes one counted bucket of errors, which is why the order of fixes comes from the taxonomy, not intuition.

What you need to know

Going from 78% to 95% means removing about three quarters of the errors. You cannot do that by guessing which sentence of the prompt to change. You need to know what the errors are.

Build the taxonomy

  1. Freeze a test set — 200 labelled documents that never change during the work.
  2. Run and collect — every field that does not match the label is one failure.
  3. Tag each failure — by cause, not by field.
  4. Count — sort the buckets by size and fix the biggest first.

A typical result for invoices:

BucketShare of failuresTargeted fix
Format (date as "3rd March", amount with commas)30%Schema-constrained output with typed fields
Invented value when field is absent25%Allow null; add an example where the field is missing
Two similar fields confused (invoice date vs due date)20%Clear field descriptions; a separate focused call
Value from the wrong section (billing vs shipping address)15%Few-shot examples from these failures
Genuinely ambiguous document10%Flag for human review

The fixes

  • Constrain the output. Structured-output modes restrict generation to your JSON Schema, so format errors become impossible rather than unlikely. Types such as date and Decimal remove a whole bucket.
  • Few-shot from the errors. Choose 5 to 8 examples that come from the failing buckets, including one where a field is truly absent and the correct output is null.
  • Decompose. If two fields are always confused, extract them in their own short call with a focused instruction. Two cheap calls often beat one clever prompt.
  • Verify high-value fields. A second check confirms each extracted value appears in the source text. A value that cannot be found is flagged, not trusted.
Python
class Invoice(BaseModel):    invoice_number: str    invoice_date: date = Field(description="Date the invoice was issued, not the due date")    due_date: date | None = Field(description="Payment due date; null if not printed")    total_amount: Decimal    gstin: str | None = Nonedef verify(inv: Invoice, source: str) -> list[str]:    return [f for f in ("invoice_number", "gstin")            if getattr(inv, f) and getattr(inv, f) not in source]

The field descriptions are sent to the model, so they carry the disambiguation. verify returns fields whose value does not appear literally in the document.

Measure per field

A 94% average can hide one field at 60%. Report accuracy per field, and set the 95% target per field where the business needs it.

A real-life example

Scenario (illustrative numbers). An accounts-payable team extracts 14 fields from supplier invoices. Average field accuracy is 78%. The engineer labels all 616 failing fields in the 200-document test set.

The biggest bucket is date format (190 failures), and structured output removes almost all of it: 84%. Adding null for absent fields and an example of a missing due date fixes most invented values: 89%. Splitting invoice date and due date into a focused call: 92%. Few-shot examples from the billing-versus-shipping failures, plus a verifier on GSTIN and invoice number: 95.4% average, with every field above 93%. The remaining failures are mostly blurry scans, which now go to a review queue.

Follow-up questions to expect

  • "How do you avoid overfitting the prompt to the test set?" — Keep a second held-out set you check only at the end, and refresh it with new documents each month.
  • "When would you move to fine-tuning?" — When a bucket survives all prompt fixes. The labelled failures are already the start of a training set.
  • "Does a bigger model solve it?" — Sometimes partly, but it costs more on every call; measure the per-bucket gain before paying for it.