Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is Automatic Prompt Engineering (APE) used for, and how do you judge the best prompt?


Three finalists on 120 held-out products3.996%6104.398%1,4504.299%720Judge scoreFacts supportedTokensCurrentCandidate ACandidate BA's lead on the judge score was inside run-to-run noise; its cost was not.
Decide the rule before you see the scores, or the biggest number wins even when it is noise.

What you need to know

What it is used for

  • Beating a hand-written baseline on a high-volume task.
  • Model migration — a prompt tuned for one model may be too prescriptive or too loose for the next. Re-optimising is faster than rewriting by hand.
  • Compression — find a shorter prompt with the same score, cutting cost on every call.
  • Continuous improvement — add new production failures to the dev set and re-run.
  • Choosing few-shot examples — optimisers such as DSPy's also pick which examples to include.

How to judge the best prompt

CheckWhy
Same fixed eval set for all candidatesOtherwise scores are not comparable
Task metric (accuracy, F1, per-field match, pass rate)Measures what the business needs
Rules checks (schema valid, length, banned words)Cheap, exact, catches regressions
LLM judge with a rubric, validated against humansFor open-ended quality
Cost and latency per callA small gain may not pay for 3× the tokens
Held-out test setGuards against overfitting to the dev set
Robustness set (hard, adversarial, other languages)Guards against brittle wins

Is the difference real?

With 200 test cases, one case is half a percentage point. A 1-point gain can come from luck, especially when the model's output varies between runs. Run each finalist 3 times and compare averages, or use a paired comparison on the same cases. Prefer the simpler, cheaper prompt when scores are within noise.

Composite decisions

Write the decision rule before you look at results, for example: "highest judge score, provided schema validity is 100%, no banned words, and cost is within 20% of the current prompt."

A real-life example

An e-commerce team optimises its product-description prompt and ends up with three finalists, each run 3 times on 120 held-out products:

PromptJudge score (1–5)Facts supportedBanned wordsTokens per call
Current3.996%2610
Candidate A4.398%01,450
Candidate B4.299%0720

Candidate A has the best judge score, but the gap with B is within the variation between runs, and it costs twice as many tokens across 50,000 descriptions a month. B wins on the pre-agreed rule. Before shipping, a copywriter reads 30 of B's outputs to check the judge — the judge had rated one overly salesy description highly, so they add "no superlatives" to the judge rubric for next time.

Follow-up questions to expect

  • "Why not just choose the highest score?" — Because small gaps can be noise, and cost, latency and safety also matter. Decide the rule in advance.
  • "How do you know the judge is right?" — Compare its scores with human ratings on a sample and fix the rubric where they disagree.
  • "How often would you re-optimise?" — When the model changes, when production failures pile up, or when the input mix shifts.