Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

How do you measure whether a prompt is effective?


The regression gate a prompt change must passFixed eval set:real cases andpast failuresRun old andnew prompt onthe same casesScore perfield orper classShip only if nothingimportant regressesThe new invoice prompt raised GSTIN accuracy and quietly dropped the total by 7 points.
An average can rise while the one field the business cares about falls, so compare per field, not per prompt.

What you need to know

"It looked good on three examples" is not evidence. A prompt change that fixes one case often breaks two others. You only see that with a fixed set of cases scored the same way each time.

  1. Collect inputs — 50 to 300 real examples, sampled from production, plus known hard cases and past failures.
  2. Define "good" — an expected label, expected fields, or a written rubric for free text.
  3. Pick metrics — the table below.
  4. Automate checks — format, schema, length, forbidden content: cheap and exact.
  5. Judge the rest — an LLM judge with a rubric, validated against humans on a sample.
  6. Compare and gate — run old and new prompts on the same set; ship only if the score rises and nothing important regresses.

Metrics by task

TaskMetric
ClassificationAccuracy, per-class precision and recall, F1, confusion matrix
ExtractionExact match per field, precision and recall per field
JSON outputPercent that parse and match the schema
RAG answersGroundedness (claims supported by the source), answer relevance
Free textRubric score from an LLM judge, plus human spot checks
CodePercent that pass unit tests

Always add cost (tokens per call) and latency (p50 and p95). A prompt that adds 2,000 tokens for one point of accuracy may not be worth it.

LLM-as-judge, done properly

A second model scores outputs against a written rubric ("1 if every claim appears in the source, else 0"). Binary or 1–3 scales are more stable than 1–10. Before trusting it, compare its scores with human labels on 50 to 100 items; if agreement is low, fix the rubric.

Python
def evaluate(prompt_version, cases, classify):    correct = 0    for case in cases:        predicted = classify(prompt_version, case["text"])        correct += predicted == case["expected"]    return correct / len(cases)# v2 must beat v1 on the same cases before it shipsprint(evaluate("v1", cases, classify), evaluate("v2", cases, classify))

The loop is trivial on purpose: the value is in the fixed cases list, not in the code. Run it in CI whenever the prompt file changes.

A real-life example

An invoice-extraction team changes its prompt to handle GST invoices from new suppliers. On the 12 new invoices it looked perfect. The eval set of 400 invoices tells a different story:

Fieldv1v2
Invoice number98%99%
GSTIN91%97%
Total amount96%89%
JSON valid100%100%

The new wording made the model pick the pre-tax subtotal as the total on multi-page invoices. The team adds "the total is the final payable amount including tax" and re-runs: total amount returns to 97% with the GSTIN gain kept. Without the eval set, a 7-point drop on the most important field would have shipped.

Follow-up questions to expect

  • "How big should the eval set be?" — Big enough that one case is not a large share of the score; 100 to 300 cases is common to start. Add every production failure you fix.
  • "How do you avoid overfitting the prompt to the eval set?" — Keep a held-out set you do not look at while tuning, and refresh it with new production samples.
  • "How do you evaluate after launch?" — Sample live traffic, run the same automated checks and judge, and watch user signals such as edits, escalations or thumbs-down.