Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
How can prompt engineering help extract information from text?
What you need to know
What makes it work
- Field definitions, not field names. "
liability_cap: the maximum amount one party can be required to pay under the contract, as stated; not insurance amounts." - Structured output so code can consume it.
- Grounding — "only what the text states; null if absent; do not infer from typical contracts."
- Evidence spans — the exact quote supporting each field. Code can check the quote exists in the document, and reviewers can verify in seconds.
- Normalisation — ISO dates, numbers without symbols, canonical names ("Govt. of India" becomes "Government of India").
- Examples for awkward cases — two values, conditional values ("90 days, or 30 days for breach").
Long documents
For a 60-page contract:
- Chunk by section or page, with a little overlap.
- Extract per chunk with evidence quotes.
- Verify each quote appears word for word in the source.
- Merge — combine chunks, remove duplicates, flag conflicts (two different governing laws).
- Review — fields with no quote or conflicting values go to a person.
Current models with long context windows can often read the whole contract at once. That saves the merge step, but evidence quotes and verification still matter.
Measuring it
For each field, on a labelled set:
precision = correct extracted values / all extracted valuesrecall = correct extracted values / all values actually presentPrecision tells you how often a filled field is wrong; recall tells you how often the model misses something. Fields with low precision need review or better definitions.
A real-life example
A legal team at a logistics company reviews 1,200 vendor contracts before renewal. They need five fields per contract: parties, governing law, termination notice, liability cap and auto-renewal.
For each field, return {"value": ..., "quote": "..."} or null.- termination_notice_days: days of notice needed to end the contract for convenience (not for breach). Number only.- auto_renewal: true only if the contract renews without action.Copy the quote word for word from <contract>.On 100 labelled contracts, per-field results are:
| Field | Precision | Recall |
|---|---|---|
| Parties | 99% | 99% |
| Governing law | 98% | 96% |
| Termination notice | 94% | 88% |
| Liability cap | 91% | 83% |
| Auto-renewal | 97% | 95% |
Liability cap is weakest, because caps are often defined across two clauses. The team adds an example of a split cap and routes every liability-cap value to a lawyer for a 30-second check against the quote. Review time falls from about 40 minutes per contract to 8.
Follow-up questions to expect
- "How do you get a confidence score?" — Use evidence (no quote means low confidence), agreement across two runs or two models, and validation rules; the model's own stated confidence is unreliable.
- "LLM or a traditional NER model?" — A trained NER model is cheaper for fixed, high-volume entity types; LLMs win when fields need definitions and judgement or change often.
- "How do you handle conflicting values?" — Keep both with their quotes and flag the record for review rather than picking one silently.