Course Content
Prompt Engineering for LLMs
3 sections · 8 lessons
Prompt Optimization & Refinement
An engineer spends a morning improving a classifier prompt. She tries a wording, pastes in three test emails, reads the outputs, likes them. Tries another wording, pastes the same three emails, likes those more. Ships it. Two weeks later the ops team reports that urgent tickets are being missed, and a review of the logs shows the new prompt is worse than the old one on the cases that matter — it just happened to look better on the three emails she kept in her clipboard.
Nothing about her process was lazy. She read every output carefully. The flaw was structural: she was measuring on three items, chosen by convenience, with no baseline, and judging by whether the output pleased her rather than whether it was right.
Three items cannot distinguish a prompt that is 70% accurate from one that is 90% accurate. Both will get all three right a good fraction of the time. What she was actually observing was noise, and she optimised towards it.
The gap between prompt engineering as a hobby and prompt engineering as a discipline is entirely here. Not in knowing more techniques — in being able to tell whether a change helped.
Build the ruler before you cut anything
You cannot improve what you cannot measure, and for prompts the measuring instrument is an evaluation set: a fixed collection of inputs with known correct outputs, which you run every candidate prompt against.
Four properties make an eval set useful.
| Property | What it means | What breaks without it |
|---|---|---|
| Real | Sampled from actual traffic or actual documents | You tune for inputs that never occur and miss the messy ones that do |
| Labelled by a human | Someone who owns the decision wrote down the right answer | You measure agreement with a guess, not correctness |
| Big enough | Typically 100–300 items for a classification task | Random variation swamps real differences |
| Covers the hard cases | Deliberately includes the confusable and the malformed | Your score is dominated by easy items and hides the failures |
The last one deserves emphasis. If 80% of real traffic is trivially easy, an eval set that mirrors traffic exactly will be 80% easy, and a prompt that fails every hard case still scores 80%. Deliberately over-sample the difficult region — then, when you report a number, be clear that it is accuracy on a hard-weighted set, not on live traffic.
Also hold some data back. Keep a second set that you look at rarely — perhaps three or four times over the life of the project. If you iterate 40 times against one set, you will end up fitting the wording of your prompt to the quirks of those particular items, and your reported number will be optimistic. A held-out set is how you find out by how much.
The first hour of prompt optimisation should produce no prompt at all. It should produce fifty labelled examples and a script that scores a prompt against them.
Metrics that mean something
"It looks better" is not a metric. Pick numbers matching what actually matters for the task.
| Metric | How to compute it | Use when |
|---|---|---|
| Exact match | Output equals expected, after trimming | Classification, extraction with a single right answer |
| Format validity | Fraction that parse and satisfy the schema | Any output a program consumes — check this first |
| Per-class recall | Of the true class-C items, how many were labelled C | Imbalanced classes; when misses on one class are costly |
| Per-class precision | Of the items labelled C, how many really were C | When false alarms are costly (paging, escalation) |
| Constraint compliance | Fraction meeting a stated rule (word count, no forbidden terms) | Generation tasks with hard requirements |
| Tokens per call | Mean input + output tokens | Always — it is the bill |
| Latency, p50 and p95 | Wall-clock per call | Anything a user waits for |
Why a single accuracy figure lies
Take an urgency classifier over 200 emails with this true distribution: 8 at level 4 (revenue stopped), 22 at level 3, 90 at level 2, 80 at level 1.
A prompt that never predicts level 4 at all, and is otherwise decent, might score 84% overall. Perfectly respectable number. But it misses all 8 of the emergencies, which are the only reason the system exists. Overall accuracy hid that completely, because the 8 items are 4% of the set and 4 points of accuracy is invisible next to the noise.
Break it out per class and the failure is unmissable:
| True level | Count | Correctly labelled | Recall |
|---|---|---|---|
| 4 — revenue stopped | 8 | 0 | 0% |
| 3 — blocked, no workaround | 22 | 15 | 68% |
| 2 — degraded | 90 | 81 | 90% |
| 1 — no time pressure | 80 | 72 | 90% |
| Overall | 200 | 168 | 84% |
Always compute the metric per class, and weight it by what an error costs. A prompt that is 84% accurate overall and 0% on the class you built the system for is a failed prompt.
Is that difference real?
This is the part that most teams skip, and it is the reason so much prompt tuning is unknowingly circular.
You test prompt A on 100 items and get 82. You test prompt B and get 86. Is B better?
Treat each item as a coin flip that lands correct with probability p. The standard error of the measured proportion is
With p=0.84 and n=100:
That is 3.7 percentage points for one standard error, so a rough 95% interval spans about ±7.3 points. Your 82 and your 86 sit comfortably inside each other's uncertainty. You have learned essentially nothing. Raise the set to 400 items and the standard error halves to about 1.8 points — the error shrinks with n, so quadrupling the data halves the noise.
Compare on the same items, and only count the disagreements
There is a much cheaper way to gain sensitivity: run both prompts on the same items and ignore every item where they agree. Items both get right, or both get wrong, carry no information about which is better.
Suppose A and B agree on 176 of 200 items. On the remaining 24: B is right and A wrong on 17, A is right and B wrong on 7. If the two prompts were genuinely equivalent, each disagreement would break either way with probability one half, so you would expect about 12 and 12. Getting 17–7 has a two-sided probability of roughly 0.06 under that assumption — suggestive, not conclusive, and enough to justify collecting more data rather than declaring victory.
1from math import comb23def paired_test(results_a: list[bool], results_b: list[bool]) -> dict:4 """Compare two prompts scored on the SAME evaluation items."""5 b_only = sum(1 for a, b in zip(results_a, results_b) if b and not a)6 a_only = sum(1 for a, b in zip(results_a, results_b) if a and not b)7 n = b_only + a_only8 if n == 0:9 return {"verdict": "identical on every item"}10 k = max(b_only, a_only)11 # Two-sided probability of a split this lopsided if the prompts tie.12 p = 2 * sum(comb(n, i) for i in range(k, n + 1)) / 2 ** n13 return {14 "b_better_on": b_only,15 "a_better_on": a_only,16 "p_value": min(p, 1.0),17 "verdict": "real difference" if p < 0.05 else "not distinguishable",18 }Two rules fall out of this and they are worth pinning up.
- Change one thing at a time. If you rewrite the instruction, add three examples and switch to JSON output in one go, a two-point gain tells you nothing about which change caused it — and one of the three may be actively harmful, hidden by the other two.
- Hold sampling fixed while comparing. Use the same model and the same sampling settings for both prompts, and the lowest temperature the model allows. Otherwise you are measuring sampling noise on top of everything else, and the same prompt will score differently on consecutive runs. Two cautions: even
temperature=0is not fully deterministic on hosted models, and many current reasoning models reject atemperaturesetting altogether. On those, run each item two or three times and count an item as correct only if the runs agree, or simply use a bigger eval set.
The loop
Optimisation is a cycle, and the diagnostic step is the one people skip.
1. RUN current prompt over the whole eval set, record every output2. TRIAGE group the failures by cause, not by input3. FIX change ONE thing aimed at the largest group4. MEASURE rerun; compare paired against the previous version5. DECIDE keep it, revert it, or gather more data6. LOG version, change made, scores, decisionStep 2 is where the value is. Reading forty failures individually produces forty small hunches. Grouping them by cause produces two or three fixes.
| Failure category | Symptom | The fix that works |
|---|---|---|
| Format | Right content, unparseable shape; markdown fences; preamble | Explicit output contract; one or two examples; a "no commentary" line |
| Vocabulary | Emits labels outside your set ("mixed", "urgent-ish") | Enumerate the allowed values in the prompt; validate in code |
| Boundary | Consistently confuses one specific pair of classes | Define the distinction by an observable test; add examples on that border |
| Missing rule | Fails a case your prompt never mentioned | State the rule; check it does not contradict existing examples |
| Reasoning | Multi-step items wrong, single-step items fine | Reasoning tokens before the answer, or split into stages |
| Ambiguity | Two of your own labellers disagree on the item | Not a prompt bug — fix the policy, then relabel |
That last row is the one that saves the most time. If your own team cannot agree on an item's label, no prompt can be right about it, and hours spent rewording will move the score randomly. Find these first and either resolve the policy or remove them from the eval set.
A worked run, start to finish
Email urgency, 200 labelled items, four levels, measured with a paired comparison at fixed sampling settings.
v1 — the naive prompt
Rate how urgent this email is.{email}Exact match 51%. Format validity 68% — the rest returned prose, or words like "fairly urgent", or a 1–5 rating when the scale has four levels. Triage: format and vocabulary dominate. Accuracy is not even meaningfully measurable yet.
v2 — define the scale and the output
Rate the urgency of the email below.4 = revenue or safety is stopped right now3 = a customer is blocked with no workaround2 = degraded but usable, or a workaround exists1 = question, request or feedback with no time pressureEmail:"""{email}"""Reply with the single digit only. No other text.Exact match 74%. Format validity 100%. Two changes, both aimed at the identified categories: the digit-only contract killed the format failures, and the anchored scale killed the vocabulary failures.
v3 — attack the specific confusion
The v2 failure breakdown showed one cause behind the biggest share of the 52 errors. Thirteen of the 22 level-3 emails were labelled 2: one customer was blocked, but the tone was calm. A second group went the other way, with angry emails about cosmetic problems rated too high. The prompt's wording invited the model to rate by how upset the sender sounded rather than by impact.
... (scale as above) ...Judge by impact, not by tone. A politely worded email describing ablocked customer is level 3. An angry email about a cosmetic issueis level 1.Email: "The invoice PDF has our old logo on it." -> 1Email: "Nobody on our team can submit timesheets since the update. We have payroll on Friday." -> 3Email: "Checkout returns 500 for all customers." -> 4Email: "Export is slow but the CSV download still works." -> 2Email:"""{email}"""Reply with the single digit only. No other text.Exact match 86%. Level-3 recall rose from 41% to 82%. Paired against v2: v3 better on 29 items, worse on 5, which is decisive.
v4 — reasoning before the digit
... same as v3, plus ...First write one sentence identifying who is blocked and from doingwhat. Then, on a new line, write:LEVEL: <digit>| Version | Exact match | Level-4 recall | Mean tokens | p95 latency |
|---|---|---|---|---|
| v1 | 51% | 25% | 210 | 0.9 s |
| v2 | 74% | 63% | 260 | 0.9 s |
| v3 | 86% | 88% | 420 | 1.0 s |
| v4 | 88% | 100% | 460 in, 45 out | 2.8 s |
Now the honest reading of v4. Overall it gained two points, and a paired comparison gave 12 items better, 8 worse — noise. But level-4 recall went from 7 of 8 to 8 of 8, and a missed level-4 email is the expensive failure this system exists to prevent. One item is far too little evidence to conclude anything, so the correct decision is not "ship v4" but "collect 40 more level-4 examples and re-test that class specifically".
Meanwhile v3 versus v2 is a clear, large, cheap win. Ship v3 now; keep the v4 question open.
The final decision is rarely "which prompt is most accurate". It is "which prompt is most accurate on the errors that cost us, at a latency and price we can pay".
Version prompts like code
A prompt that is edited in place, in a string literal, by whoever is on shift, will regress and nobody will know when. Give it the same treatment as any other artefact your system depends on.
1from dataclasses import dataclass23@dataclass(frozen=True)4class Prompt:5 id: str6 version: int7 template: str8 allowed_outputs: frozenset[str]9 eval_score: float # on the eval set named below10 eval_set: str # e.g. "urgency-v2-200items"11 notes: str # what changed and why1213URGENCY_V3 = Prompt(14 id="email-urgency",15 version=3,16 template=URGENCY_V3_TEMPLATE,17 allowed_outputs=frozenset("1234"),18 eval_score=0.86,19 eval_set="urgency-v2-200items",20 notes="Added impact-not-tone rule and 4 boundary examples. "21 "Fixed level-3 under-rating. Paired vs v2: 29 better, 5 worse.",22)Three details matter more than the exact structure. The allowed outputs live beside the template, so the validator and the prompt cannot drift apart. The score is stored with the name of the eval set it was measured on, because a score without its ruler is meaningless. And the note records the paired result, so six months later nobody has to re-derive whether the change was real.
Log the prompt version alongside every production call. When quality changes, the first question is always "what changed?", and without the version stamped on each response you cannot answer it.
Where optimisation goes wrong
| Mistake | Why it is costly |
|---|---|
| Judging by reading a handful of outputs | Cannot detect differences smaller than about 20 points; you optimise towards noise |
| Changing several things at once | You cannot attribute the result, and a harmful change hides behind a helpful one |
| Tuning against the same 50 items for weeks | The prompt fits those items; live performance is lower and you do not know by how much |
| Adding text and never removing any | Prompts bloat to 2,000 tokens of which a third contradicts the rest |
| Ignoring cost and latency until launch | The winning prompt turns out to be unaffordable at real volume |
| Treating disagreement between labellers as a model failure | Days spent rewording a prompt for items that have no correct answer |
The fourth is worth a note because it creeps up on you. Every fix adds a line, nothing is ever deleted, and after two months the prompt contains a rule from v2 that a v6 example directly contradicts. Periodically try removing lines and re-measuring. Anything you can delete without a score change was costing you money on every call — and might have been quietly hurting.
When you build something with this
The discipline compresses into a short sequence, and the order is not negotiable.
- Labelled data before prompt work. A hundred real items with human labels, plus a script that scores a prompt in one command. Until that exists, every improvement you make is unverifiable.
- Fix format before accuracy. Unparseable output makes accuracy unmeasurable, and format problems are the cheapest to fix. Get validity to 100%, then start on the judgements.
- Group failures by cause, then fix the biggest group. One change per iteration, targeted at a category you have actually counted.
- Compare paired, at fixed sampling, and respect the noise floor. On 200 items, differences under about four points are not evidence. Either gather more data or accept that the two prompts are the same.
- Weight by cost of error, not by count. Report per-class numbers and decide on the class that matters.
- Record version, change, score and eval-set name — every time. Six months from now, the log is the only thing standing between you and repeating the whole exercise.
None of this is glamorous, and it is the entire difference between a prompt that survives contact with real traffic and one that merely looked good on the three examples somebody had open.