Course Content
Prompt Engineering for LLMs
3 sections · 8 lessons
Mini-Project: Build a Prompt Library
Search a six-month-old codebase for the string "Summarise" and count what you find. In most teams the answer is depressing: four variants of essentially the same prompt, in four files, written by three people. One of them says "in 3 bullet points", another says "briefly", a third has a stray instruction about JSON left over from a different feature. Nobody knows which performs best because nobody has ever compared them. When the model provider ships a new version, nobody knows which of the four still works.
Meanwhile the good prompt — the one someone spent a day tuning and which really does work — lives in a Slack message from March.
This is what happens when prompts are treated as strings rather than as artefacts. They are neither tested, versioned, owned, nor discoverable, and so every improvement anyone makes evaporates. A prompt library is the fix: a small, boring piece of infrastructure that makes a prompt a thing with a name, a version, a measured score, and a test that runs.
What you are going to build is three working prompts and the machinery around them. The three prompts matter less than the machinery, because the machinery is what makes the fourth, fifth and fiftieth prompt cheap.
What you are building
Four components. None is large.
| Component | What it is | Why it exists |
|---|---|---|
| Registry | A structured record per prompt: id, version, template, variables, allowed outputs, eval set name, score | So a prompt has an identity that survives being edited |
| Eval sets | A file of labelled input/expected pairs per prompt | So "better" is a measurement, not an opinion |
| Harness | One command that scores any prompt version against its eval set | So testing a change costs thirty seconds, not an afternoon |
| Validators | Per-prompt code that checks the output is usable | Because a prompt instruction lowers the probability of bad output; it does not prevent it |
The rule that makes a prompt library work, and the one everyone is tempted to break: a prompt without an eval set does not go in the library. It can live in a scratch file until it has one.
The registry entry
1from dataclasses import dataclass, field2from typing import Callable34@dataclass(frozen=True)5class Prompt:6 id: str # "support.sentiment"7 version: int8 template: str # uses {named} placeholders9 variables: tuple[str, ...] # what must be supplied10 eval_set: str # filename of the labelled data11 score: float # exact-match on that eval set12 validator: Callable[[str], object | None] # None = unusable13 changelog: str # what changed and the paired result1415 def render(self, **kwargs) -> str:16 missing = set(self.variables) - set(kwargs)17 if missing:18 raise ValueError(f"{self.id} missing: {sorted(missing)}")19 return self.template.format(**kwargs)2021REGISTRY: dict[str, Prompt] = {}2223def register(p: Prompt) -> Prompt:24 key = f"{p.id}@{p.version}"25 if key in REGISTRY:26 raise ValueError(f"{key} already registered - bump the version")27 REGISTRY[key] = p28 REGISTRY[p.id] = p # bare id points at the newest29 return pThree details are doing real work. render raises on a missing variable instead of leaving a literal {ticket} in the prompt — a bug that otherwise produces confident nonsense and no error. The validator travels with the template, so the allowed outputs in the prompt and the allowed outputs in the checking code cannot drift apart. And registering a duplicate version fails loudly, which stops the most common library failure: someone editing a template in place, leaving the version number and the score untouched, and silently invalidating both.
The harness
1import json, time2from collections import Counter34def evaluate(prompt: Prompt, call, eval_dir: str = "evals",5 sampling: dict | None = None) -> dict:6 # e.g. {"temperature": 0} for models that accept it; {} for7 # reasoning models that reject sampling parameters.8 sampling = sampling or {}9 items = [json.loads(l) for l in open(f"{eval_dir}/{prompt.eval_set}")]10 correct = invalid = 011 latencies, confusions = [], Counter()1213 for item in items:14 t0 = time.perf_counter()15 raw = call(prompt.render(**item["input"]), **sampling)16 latencies.append(time.perf_counter() - t0)1718 parsed = prompt.validator(raw)19 if parsed is None:20 invalid += 121 continue22 if parsed == item["expected"]:23 correct += 124 else:25 confusions[(item["expected"], parsed)] += 12627 n = len(items)28 latencies.sort()29 return {30 "prompt": f"{prompt.id}@{prompt.version}",31 "n": n,32 "exact_match": correct / n,33 "format_validity": (n - invalid) / n,34 "p95_latency": latencies[int(0.95 * n) - 1],35 "top_confusions": confusions.most_common(5),36 }The sampling argument is deliberate. While comparing two prompts you want the sampling settings identical and the noise as low as the model allows, otherwise you are measuring randomness on top of the difference you are trying to detect. It is a parameter, not a hard-coded temperature=0, because many current reasoning models reject a temperature setting; keep the model name and its sampling settings together in configuration. The top_confusions field is the one you will actually use — it turns a disappointing score into a named problem within seconds.
Prompt one: sentiment with a usable middle
Job: label product reviews so the product team can filter by sentiment.
Start deliberately naive, so you can see what fails rather than guessing.
v1:What is the sentiment of this review?{review}Run the harness over 60 labelled reviews. Expect something in the region of 55–70% exact match and format validity well under 100%, with outputs like "Mostly positive", "Positive sentiment overall", "Mixed — the customer liked X but disliked Y."
Read the failures and you will find the real problem is not the model. It is that your own labelling policy is undefined for mixed reviews. Half your eval set probably disagrees with itself. Fix the policy first — a prompt cannot be correct about a question you have not answered.
v2:Classify the sentiment of the product review below. positive - the reviewer would recommend the product negative - the reviewer would advise against it neutral - the reviewer describes it without a recommendation either way, or the review has no opinion contentA review that contains both praise and complaint is classified bythe reviewer's overall recommendation, not by the balance of words.Review:<review>{review}</review>Reply with exactly one word: positive, negative or neutral.Why this works. Each label is anchored to a single observable question — would this person recommend it? — which two labellers can apply consistently, and which the model can check against the text. The mixed-review rule removes the ambiguity that was producing most of the disagreement. The closed word list makes anything outside the set an improbable continuation, and the validator catches the rest.
1SENTIMENT_LABELS = {"positive", "negative", "neutral"}23def sentiment_validator(raw: str):4 word = raw.strip().strip(".").lower()5 return word if word in SENTIMENT_LABELS else NoneIf v2 still confuses one specific pair — typically neutral and positive — add three or four examples drawn from the confusions the harness reported, not from imagination. Real boundary cases teach the boundary; invented ones teach what you wish the boundary were.
Record each version's score in the changelog. You are building the habit that makes the library valuable: no version enters without a number attached.
Prompt two: summarising technical documentation
Job: turn a page of API documentation into a summary a developer can scan in fifteen seconds.
v1:Summarise this documentation.{doc}The output will be fluent and largely useless. Two things go wrong, and both are predictable.
First, compression is undefined. The prompt never says how short, and it never says what the summary is for, so the model cannot tell which details are load-bearing. A summary is only defined relative to a purpose.
Second, and more seriously, invention is free. The model is generating plausible text about an API. Plausible text about an API includes parameter names, default values and rate limits — whether or not the source document contains them. You will see a confident sentence about a default timeout that appears nowhere in the input.
v2:Summarise the API documentation below for a developer decidingwhether this endpoint does what they need.Rules:- Every statement must be supported by the documentation. Do not state a parameter, default, limit or error code that is not written there.- If the documentation does not state something a section asks for, write "not documented". Do not infer it from the endpoint name or from how similar APIs usually behave.- Copy identifiers exactly, including case and underscores.Output exactly:DOES: one sentence, what the endpoint doesAUTH: what authentication is required, or "not documented"REQUIRED: required parameters as a comma-separated list, or "none"LIMITS: rate limits or size limits, or "not documented"GOTCHA: the one behaviour most likely to surprise someone, or "nothing unusual documented"Documentation:<doc>{doc}</doc>Why this works. The named reader and their decision define what matters, replacing "important" with a criterion. The fixed sections replace an undefined compression ratio with a budget. Most importantly, "not documented" gives absence a legal way to be expressed — without it, a model asked for a LIMITS line with no limits in the source will produce a plausible limit, because producing something is the overwhelmingly more probable continuation. "Do not infer it from how similar APIs usually behave" targets the specific mechanism by which that invention happens.
The validator here is not exact-match; the useful check is groundedness.
1import re23def doc_summary_validator(raw: str):4 required = ["DOES:", "AUTH:", "REQUIRED:", "LIMITS:", "GOTCHA:"]5 if not all(s in raw for s in required):6 return None7 return raw89def ungrounded_identifiers(summary: str, source: str) -> list[str]:10 """Identifiers in the summary that never appear in the source."""11 ids = set(re.findall(r"\b[a-z]+_[a-z_]+\b", summary))12 return sorted(i for i in ids if i not in source)Score this prompt on two numbers: structure validity (all five sections present) and groundedness (no invented identifiers or figures). Exact-match scoring does not apply to open-ended generation, so pick metrics that match the actual failure mode rather than forcing the task into a shape your harness already handles.
Absence has to be given a legal way to be expressed. A section header with no matching content in the source is a strong pull towards invention, and "not documented" is the release valve.
Prompt three: drafting a support reply
Job: draft a reply to a customer complaint for an agent to review and send.
This one is different in kind. There is no single correct output, so you cannot score it by exact match. What you can score is constraint compliance — and it turns out that most of what makes a support reply good or bad is expressible as a constraint.
v1:Write a helpful, empathetic reply to this customer complaint.{complaint}The output will be smooth and will contain at least one of these: an apology so effusive it reads as insincere, a promise about a timeline nobody authorised, an admission of fault with liability implications, or a paragraph of filler before anything useful. Every one of these is the model reproducing the centre of mass of customer-service writing, which is exactly where you do not want to be.
v2:Draft a reply to the customer complaint below. A support agent willreview and edit it before it is sent.Facts you may use:{known_facts}Structure, in this order:1. Acknowledge the specific problem in the customer's own terms. One sentence. Do not open with "Thank you for reaching out."2. State what is actually true about the cause, using only the facts above. If the cause is unknown, say it is being investigated.3. State the single next action and who takes it.4. Close with one sentence. No offer of compensation.Constraints:- 90-130 words.- No timeline unless a specific date appears in the facts above.- Never say "we apologise for any inconvenience".- Never accept fault or use the words "our error", "our fault", "we failed". Describe what happened, not who is to blame.- If the facts above do not cover what the customer is asking, end the draft with the single line: ESCALATE: <what the agent needs to find out>Complaint:<complaint>{complaint}</complaint>Why this works. The known_facts variable is the most important part of the design: the model can only invent causes and timelines if you leave it a gap to fill, so the prompt supplies the facts and forbids anything outside them. Banning specific phrases is far more effective than asking for sincerity, because the phrases are the actual high-probability tokens producing the problem. The word range is checkable. The fault-language prohibition is a real business requirement encoded as a constraint rather than trusted to judgement. And the ESCALATE line converts the dangerous case — a complaint the facts do not cover — from an invented answer into a routed one.
1BANNED = ["apologise for any inconvenience", "our error", "our fault",2 "we failed", "thank you for reaching out"]34def reply_validator(raw: str):5 text = raw.strip()6 lower = text.lower()7 body = text.split("ESCALATE:")[0] # the ESCALATE line is not counted8 if not 90 <= len(body.split()) <= 130:9 return None10 if any(b in lower for b in BANNED):11 return None12 return textThe eval set for this prompt is not input/expected pairs. It is a set of complaints plus the facts available for each, and the score is the fraction of drafts passing every constraint. That is a perfectly respectable metric, and it catches regressions just as reliably as exact match does.
Documenting an entry so it is still usable in a year
Each prompt in the library needs a short record. Not a document — five fields somebody will actually maintain.
id: support.reply_draftversion: 3purpose: Draft a customer reply for agent review. Never sent unedited.inputs: complaint (untrusted user text), known_facts (from the incident record; may be empty)output: 90-130 word draft, optionally followed by ESCALATE: line Validator: reply_validator (word count + banned phrases)eval: evals/reply_draft_40.jsonl - 40 real complaints with their incident facts. Metric: constraint compliance. v1 41% v2 88% v3 94%known gaps: Fails on complaints covering two unrelated issues - drafts address only the first. 3 of 40 eval items. Not yet fixed.changed: v3 added the ESCALATE line. Paired vs v2: 8 better, 1 worse. Escalation rate on live traffic: 11%.The known gaps field is the one people leave out and the one that saves the most time. Every prompt has inputs it handles badly. Writing them down stops three different people rediscovering the same limitation, and it tells whoever picks the prompt up next where the work is.
What to measure across all three
| Prompt | Primary metric | Secondary | Failure to watch |
|---|---|---|---|
| Sentiment | Exact match against labels | Format validity | Neutral/positive confusion; mixed-review policy drift |
| Doc summary | Groundedness (no invented identifiers) | Structure validity | Inferring standard API behaviour not in the source |
| Support reply | Constraint compliance | Escalation rate | Unauthorised timelines; fault language |
Notice that the three metrics are genuinely different, and that this is not a nuisance to be smoothed over. A harness that only knows how to compute exact match will push you towards forcing every task into a classification shape, which is how teams end up with an unmeasured generation prompt sitting in production. Let the validator be per-prompt, and let the metric follow the failure mode.
Not every prompt can be scored by exact match, and that is not a reason to leave one unmeasured. Pick the metric that matches the failure mode — groundedness, structure validity, constraint compliance — and the harness works just as well.
Making the library survive
Building it takes an afternoon. Keeping it takes three habits, and libraries die when any one of them lapses.
Version on every edit, without exception. The moment someone changes a template in place and leaves the score attached, every number in the registry becomes a claim about a prompt that no longer exists. Making duplicate registration throw is worth more than any amount of policy, because it turns discipline into something the code enforces.
Re-run the harness on a schedule, not only on change. Providers update models. Your traffic shifts. A prompt scoring 89% in March can score 80% in September with nothing in your repository having changed. If the harness runs weekly in CI and posts the numbers, you find out in a week; if it only runs when someone edits a prompt, you find out from a customer.
Refresh the eval sets from live traffic. Every eval set is a snapshot of the inputs that existed when it was written. Sample fifty real inputs a month, label them, add them, and retire items that no longer represent anything. A prompt that scores 94% on data from a year ago is telling you about last year.
One last thing worth being deliberate about: when someone reports that a prompt is behaving badly, the first move is to add their case to the eval set, before changing a single word of the template. If the case does not reproduce as a failure in the harness, the problem is somewhere else — retrieval, an upstream field, a variable that arrived empty — and rewording the prompt will waste a day. If it does reproduce, you now have a test that will stop it coming back, which is the entire reason for building any of this.