Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What is Automatic Prompt Engineering (APE) used for, and how do you judge the best prompt?
What you need to know
What it is used for
- Beating a hand-written baseline on a high-volume task.
- Model migration — a prompt tuned for one model may be too prescriptive or too loose for the next. Re-optimising is faster than rewriting by hand.
- Compression — find a shorter prompt with the same score, cutting cost on every call.
- Continuous improvement — add new production failures to the dev set and re-run.
- Choosing few-shot examples — optimisers such as DSPy's also pick which examples to include.
How to judge the best prompt
| Check | Why |
|---|---|
| Same fixed eval set for all candidates | Otherwise scores are not comparable |
| Task metric (accuracy, F1, per-field match, pass rate) | Measures what the business needs |
| Rules checks (schema valid, length, banned words) | Cheap, exact, catches regressions |
| LLM judge with a rubric, validated against humans | For open-ended quality |
| Cost and latency per call | A small gain may not pay for 3× the tokens |
| Held-out test set | Guards against overfitting to the dev set |
| Robustness set (hard, adversarial, other languages) | Guards against brittle wins |
Is the difference real?
With 200 test cases, one case is half a percentage point. A 1-point gain can come from luck, especially when the model's output varies between runs. Run each finalist 3 times and compare averages, or use a paired comparison on the same cases. Prefer the simpler, cheaper prompt when scores are within noise.
Composite decisions
Write the decision rule before you look at results, for example: "highest judge score, provided schema validity is 100%, no banned words, and cost is within 20% of the current prompt."
A real-life example
An e-commerce team optimises its product-description prompt and ends up with three finalists, each run 3 times on 120 held-out products:
| Prompt | Judge score (1–5) | Facts supported | Banned words | Tokens per call |
|---|---|---|---|---|
| Current | 3.9 | 96% | 2 | 610 |
| Candidate A | 4.3 | 98% | 0 | 1,450 |
| Candidate B | 4.2 | 99% | 0 | 720 |
Candidate A has the best judge score, but the gap with B is within the variation between runs, and it costs twice as many tokens across 50,000 descriptions a month. B wins on the pre-agreed rule. Before shipping, a copywriter reads 30 of B's outputs to check the judge — the judge had rated one overly salesy description highly, so they add "no superlatives" to the judge rubric for next time.
Follow-up questions to expect
- "Why not just choose the highest score?" — Because small gaps can be noise, and cost, latency and safety also matter. Decide the rule in advance.
- "How do you know the judge is right?" — Compare its scores with human ratings on a sample and fix the rubric where they disagree.
- "How often would you re-optimise?" — When the model changes, when production failures pile up, or when the input mix shifts.