Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

What is Automatic Prompt Engineering (APE)?


Prompt writing as a searchSeed promptand 150labelled ticketsModel proposes5 candidatesScore each onthe dev setKeep thebest, revisefrom failuresConfirm on 50held-out ticketsFour rounds took the classifier from 86 to 91 percent on tickets it never saw.
The optimiser can only climb the metric you give it, so a weak metric produces a prompt that is better at nothing.

What you need to know

The loop

  1. Seed — describe the task and give input-output examples, or a starting prompt.
  2. Propose — a model writes several candidate prompts.
  3. Score — run each candidate on a development set and compute the metric.
  4. Select and vary — keep the top few and ask the model for variations or fixes based on failed cases.
  5. Stop — when the score plateaus; confirm the winner on a held-out test set.

Related methods

  • OPRO (optimisation by prompting, 2023) — shows the model past prompts with their scores and asks for a better one.
  • DSPy — a framework where you write the program's steps and a metric, and optimisers (for example MIPROv2 and GEPA) search for instructions and few-shot examples automatically. When you switch models, you re-run the optimiser instead of rewriting prompts by hand.
  • Prompt improver tools in provider consoles — rewrite a draft prompt into a more structured one. A good start, but still needs your eval set.

What it needs

  • A metric you trust. If the metric is a weak LLM judge, the optimiser learns to please the judge.
  • Enough data — dozens to a few hundred labelled examples, split into development and test.
  • A budget — every candidate is scored on every development example, so 20 candidates × 200 examples is 4,000 calls per round.

A real-life example

A support team's ticket classifier scores 86% on 200 labelled tickets with a hand-written prompt. They run an optimisation loop with 150 tickets for development and 50 held out for testing.

Text
Meta-prompt to the optimiser model:Here is a prompt for classifying support tickets, its accuracy (86%), and12 tickets it got wrong, with the correct labels. Propose 5 revised promptsthat would fix these errors without breaking the ones it got right.

After four rounds the best candidate scores 93% on development and 91% on the held-out test set. The main change was one the team had not thought of: a sentence saying that a ticket mentioning both a late order and a refund request is labelled by the root cause, delivery. The team reads the winning prompt, removes a strange sentence that did not affect the score, and ships it.

Follow-up questions to expect

  • "Does APE replace prompt engineers?" — No. Someone must define the task, build the eval set and metric, and review the winning prompt; the optimiser only searches.
  • "What is the risk?" — Overfitting to the development set or to a flawed judge. A held-out test set and human review of the final prompt control it.
  • "When is it worth the cost?" — For high-volume prompts where a few points of accuracy matter, and when switching models.