Course Content
Prompt Engineering for LLMs
3 sections · 8 lessons
Zero-Shot & Few-Shot Prompting
A support team wants incoming tickets routed automatically. They have used the same four categories for years: Billing, Account, Technical, Feedback. You write the obvious prompt.
Classify this support ticket as Billing, Account, Technical or Feedback.Ticket: "I was charged twice this month and now I can't log in."The model says Billing. Reasonable. But this company routes double-charge-plus-lockout tickets to Account, because their account team owns anything where a customer is locked out, regardless of cause. The model could not have known that. It made the choice any careful outsider would make, and it was wrong for this company.
You try to fix it with more instruction. You add a paragraph explaining that Account covers lockouts even when the trigger was a payment problem. Accuracy improves, then you hit the next case: "The export button does nothing on the reporting page and I'd like a refund for this month." Is that Technical or Billing? You add another paragraph. Then another. After an afternoon you have a 600-word rulebook that is still ambiguous in places, and the model's agreement with your team hovers around 70%.
Then you delete most of the rulebook and paste in eight real tickets from last month with the labels your team actually assigned. Agreement jumps to the high eighties. You did not explain the rule better. You stopped explaining and started demonstrating.
Understanding why that works — and precisely when it does not — is what separates prompting from guesswork.
Zero-shot: asking without showing
Zero-shot prompting means giving the model a task description and the input, with no worked examples of the task being performed. "Zero shots" means zero demonstrations.
Translate the following sentence into formal Japanese.Sentence: The meeting has been moved to Thursday.It is worth pausing on the fact that this works at all. A model whose only training objective was to predict the next token has no reason to treat "Translate the following sentence" as a command rather than as the opening of a document about translation. Base language models, before any further training, often do exactly that — you ask a question and they continue with more questions, because a list of questions is a plausible document.
The reason production models follow instructions is an additional training stage, usually called instruction tuning, in which the model is trained on many thousands of examples of the form instruction → appropriate response. That stage makes "text that looks like an instruction" strongly predictive of "text that looks like a compliant response".
Zero-shot is not a special mode. It is the model recognising your text as an instruction because it has seen an enormous number of instruction-and-response pairs, and completing the pattern.
This tells you exactly when zero-shot will be strong: when your task is one the model has seen thousands of variants of, and when the output format follows conventionally from the task.
| Zero-shot is usually enough | Zero-shot usually struggles |
|---|---|
| Translation, grammar correction, tone rewriting | Labels whose meaning is specific to your organisation |
| Summarising into prose | An exact output format with fiddly details |
| Standard sentiment (positive/negative/neutral) | Ordinal judgements — severity 1–5, priority tiers |
| Extracting obvious entities: dates, emails, names | Domain conventions: legal clause types, clinical coding |
| Answering general-knowledge questions | Consistent style across thousands of calls |
Making zero-shot as strong as it can be
Before reaching for examples, it is worth exhausting what a well-specified zero-shot prompt can do, because examples cost tokens on every single call.
BAD:Rate how urgent this email is.Email: "Hi — the payment page is showing a 500 error for allcustomers since about 9am. Nothing is going through."GOOD:Rate the urgency of the email below on this scale:4 = revenue or safety is stopped right now3 = a customer is blocked, no workaround2 = degraded but usable, or a workaround exists1 = question, request or feedback with no time pressureEmail:"""Hi — the payment page is showing a 500 error for all customerssince about 9am. Nothing is going through."""Reply with the single digit only.Why the second works. "How urgent" is a request for a judgement on an undefined scale, so the model invents a scale — sometimes 1–5, sometimes words like "high", sometimes a paragraph. Defining the scale replaces invention with selection. More importantly, each level is anchored to an observable condition rather than a feeling. "Revenue is stopped right now" is something you can check against the email's contents; "very urgent" is not. That anchoring is what makes two different runs, on two different emails, apply the same standard.
If your zero-shot prompt is failing, the first question is not "should I add examples?" It is "have I actually defined the categories, or have I named them and hoped?"
Few-shot: showing instead of telling
Few-shot prompting puts a handful of completed input/output pairs into the prompt before the real input. Two examples is two-shot, five is five-shot.
Classify each support ticket into exactly one category.Ticket: I was charged twice this month and now I can't log in.Category: AccountTicket: My invoice shows the old price even though I downgradedin March.Category: BillingTicket: The CSV export produces an empty file on Safari.Category: TechnicalTicket: The new dashboard is a huge improvement, thanks.Category: FeedbackTicket: Export button does nothing and I'd like a refund forthis month.Category: TechnicalTicket: {new_ticket}Category:The lockout rule and the refund-plus-bug rule are both in that prompt, but neither is stated. They are shown. And the model has a much easier job inferring "lockouts go to Account" from a matching case than from a paragraph of prose it must apply correctly to a novel situation.
What is actually happening — and what is not
This is the point at which language gets misleading. Few-shot prompting is often called in-context learning, and "learning" makes people imagine the model is being trained. It is not.
Nothing about the model changes. Its weights after your few-shot call are bit-for-bit identical to before. The examples do their work entirely by sitting in the context and altering which continuation is most probable.
Three concrete effects follow from that, and each one predicts a real behaviour you will observe.
1. Format lock. Once the context contains four blocks of the shape Ticket: … / Category: …, the most probable continuation after a fifth Ticket: block is a Category: line containing one short word. Prose becomes improbable — not forbidden, just outcompeted. This is why few-shot is the most reliable tool for format control, and why it works even when the format is difficult to describe in words.
2. Task location. A model's capabilities are vast and mostly latent; your prompt selects which one to run. Examples narrow that selection sharply. Given the word "Category" alone, the model must guess your taxonomy. Given four labelled cases, the label space is visible and closed.
3. Boundary transfer. The genuine payload of a good example set is the hard cases. An example that any reasonable person would label identically teaches almost nothing. The double-charge lockout example teaches something no amount of instruction was conveying.
There is a useful research finding here that is often half-remembered. In experiments on standard classification tasks (Min et al., 2022, "Rethinking the Role of Demonstrations"), replacing the correct labels in few-shot examples with random labels from the right set degrades performance far less than you would expect. The interpretation is that a large share of the benefit comes from showing the format, the label vocabulary, and the kind of input — not from the input-to-label mapping itself.
Do not over-generalise from that into "labels don't matter". They matter most in exactly the situation you are usually in: a task where your labelling convention differs from the obvious one. The honest summary is that examples deliver two separable things, and you should know which you are buying.
| What examples supply | How many you need | Signal if you are short |
|---|---|---|
| Output format and label vocabulary | 1–2 is usually enough | Output is right in substance, wrong in shape |
| Your specific decision boundaries | 3–8, covering the confusable pairs | Format is perfect, judgements disagree with your team |
Choosing the examples
Example selection is where most few-shot prompts are quietly sabotaged. Four rules, each with a mechanism behind it.
Cover the confusions, not the obvious cases
Pull real, previously-labelled items and pick the ones your own team argued about. If Technical and Billing are the pair that gets misrouted, at least two examples should sit on that border.
WEAK EXAMPLE SET (all unambiguous):"My card was declined." -> Billing"The site is down." -> Technical"Love the new design." -> FeedbackSTRONG EXAMPLE SET (sits on the borders):"Charged twice, now locked out." -> Account"Export is broken, I want a refund." -> Technical"Great product but your pricing page lies about what's included." -> FeedbackThe weak set demonstrates format and nothing else. The strong set encodes three decisions a newcomer would get wrong.
Balance the labels, and watch the last one
If six of your eight examples are labelled Technical, you have supplied a prior. The model's output distribution shifts towards the majority label, and it will over-predict Technical on genuinely ambiguous inputs. This bias is real and measurable, and it is stronger for the labels that appear near the end of your example list, because recent context has more influence on the next token than distant context.
Two practical consequences: keep the label counts roughly even unless you deliberately want to encode a base rate, and do not let your final example be from your most-common class by accident. If you can afford it, shuffle the example order across a few variants and check whether accuracy moves — if it does, your example set is too small or too imbalanced.
Keep the format byte-identical
Every example must use the same field names, the same separators, the same capitalisation, the same spacing. This is not fussiness. The whole mechanism is pattern completion, and an inconsistent pattern is a weak one.
BROKEN — three formats in three examples:Input: The app crashes on launchOutput: TechnicalTicket - "Charged twice" => BillingQ: Love the updateA: feedbackBROKEN — the case of the label also drifts (Technical / feedback),so the model has no consistent target to copy.Keep them short
Examples are paid for on every call, in both money and latency. A five-shot prompt with 400-word examples adds 2,000 tokens to every request, forever. Trim each example to the shortest version that still carries its distinction. If a 25-word ticket demonstrates the boundary as well as a 300-word one, use the 25-word one.
The real trade-off
| Zero-shot | Few-shot | |
|---|---|---|
| Prompt tokens per call | Low | High — examples repeat on every request |
| Latency | Lower | Higher (more input to process) |
| Format reliability | Moderate | High |
| Encodes house conventions | Only if you can write them out | Yes, by demonstration |
| Effort to build | Minutes | Hours — you need labelled data |
| Maintenance risk | Low | Examples go stale as your taxonomy shifts |
| Best fit | Common tasks, conventional output | Custom labels, strict formats, subtle boundaries |
Run the numbers before committing. A six-shot prompt whose examples total 900 tokens, called 200,000 times a month, spends 180 million input tokens on examples alone. If a well-specified zero-shot prompt gets you 91% and the six-shot gets 94%, whether that is worth it depends entirely on what a mistake costs you. Sometimes it obviously is. Sometimes the answer is to keep two examples instead of six and take 93%.
Examples are not free and they are not automatically better. They buy format reliability and boundary knowledge, and you pay for them on every single call.
Failure modes you should expect
The model copies an example
When a real input closely resembles one of your examples, the model sometimes returns that example's output verbatim, including details that do not apply. Make examples clearly distinct from likely inputs, and if you see verbatim reuse, add a line such as "The examples show the required format. Base your answer only on the input below them."
Examples contradict the instructions
You write "always reply with a single lowercase word", and one of your examples reads Category: Account. The model now has a demonstration that disagrees with the rule, and demonstrations generally win. Whenever you edit the instruction, re-read every example to confirm it still obeys it. This is the single most common bug in few-shot prompts, and it is invisible until you look for it.
Examples drift out of date
Six months later the team has added a Security category and moved lockouts into it. The prompt still shows lockouts as Account. Nothing errors; the model simply keeps applying last year's policy. Store your examples with the labelled dataset they came from and re-derive them whenever the taxonomy changes.
Position sensitivity masquerading as accuracy
You measure 94% on your test set. You reorder the examples and get 89%. Neither number is the "true" one — the volatility itself is the finding. It means the example set is not strong enough to determine the answer on its own, and the ordering is filling the gap. Add examples on the confusable boundaries until reordering stops moving the score much.
Selecting examples per request
A static example block is the right starting point. When a fixed set stops being enough — usually because you have many categories or a wide input distribution — the next step is choosing examples per request based on similarity to the incoming input.
1def build_prompt(new_ticket: str, pool: list[dict], embed, k: int = 4) -> str:2 """pool: labelled tickets, each {'text':..., 'label':..., 'vec':...}"""3 q = embed(new_ticket)4 ranked = sorted(pool, key=lambda ex: -cosine(q, ex["vec"]))56 # Take the nearest neighbours, but cap per label so a dominant7 # class cannot supply every example and skew the prediction.8 chosen, per_label = [], {}9 for ex in ranked:10 if per_label.get(ex["label"], 0) >= 2:11 continue12 chosen.append(ex)13 per_label[ex["label"]] = per_label.get(ex["label"], 0) + 114 if len(chosen) == k:15 break1617 shots = "\n\n".join(18 f"Ticket: {ex['text']}\nCategory: {ex['label']}" for ex in chosen19 )20 return (21 "Classify each support ticket into exactly one category.\n\n"22 f"{shots}\n\nTicket: {new_ticket}\nCategory:"23 )The per-label cap is the part worth stealing. Pure nearest-neighbour selection tends to return several examples of the same class — similar tickets usually share a label — which hands the model a lopsided prior pointing at the answer it already leans towards. Capping per label keeps the alternatives visible.
When you build something with this
Treat the choice as a sequence rather than a preference, and stop as soon as the task is solved.
- Start zero-shot, but specified properly. Enumerate the allowed outputs, define each one by an observable condition, state the output shape. A large fraction of "few-shot fixed it" stories are really "someone finally defined the categories".
- Build a small labelled test set before you tune anything. Thirty to fifty real, human-labelled items. Without it you cannot tell an improvement from a change, and you will spend days rewriting prose based on the last three outputs you happened to look at.
- Read the errors and name the failure. Wrong format means you need one or two examples. Wrong judgement means you need boundary examples on the specific pairs that are being confused. These are different problems with different fixes.
- Add examples from your labelled set, not from your imagination. Invented examples encode what you think your convention is. Real ones encode what it actually is.
- Re-measure after every change, and check the cost. Track accuracy and prompt tokens together. An extra four examples that buy half a percentage point are usually not worth paying for on every call for the next two years.
Underneath all of it sits one idea. You are not teaching the model anything; you are constructing a context in which the answer you want is the most probable thing to say next. Examples are simply the most direct way to build such a context — which is exactly why they must be consistent, balanced, current, and drawn from the boundaries where the real disagreements live.