Course Content
Prompt Engineering Mastery
6 sections · 32 lessons
What is few-shot prompting and why do we use it?
What you need to know
Few-shot prompting is also called in-context learning: the model learns the task from examples inside the prompt, without any change to its weights. Nothing is stored; the next call starts fresh.
Why examples work
Examples are the strongest signal in a prompt. A model matches their length, tone, structure and labels more faithfully than it follows an adjective like "concise". That makes them powerful for:
- Format — showing is clearer than describing.
- Your definitions — where the line between two labels sits for your business.
- Style — the rhythm of a brand voice is hard to describe but easy to show.
How to choose examples
- Diverse: different lengths, phrasings and labels, so the model learns the rule, not one surface pattern.
- Balanced: if four of five examples are "billing", the model leans towards "billing".
- Hard cases first: spend your examples on borderline inputs, not obvious ones.
- Correct: a wrong example is copied as faithfully as a right one.
- Wrapped clearly: put examples in tags such as
<example>so the model does not confuse them with the real input.
Costs and limits
Examples are paid for on every call. Five examples of 80 tokens each add 400 input tokens; at 1 million calls a month that is 400 million extra tokens. Prompt caching reduces the cost of a fixed example block, but it is still worth dropping examples that do not measurably help.
On current models, a single "gold" example can over-constrain the output — every result starts to look like it. If you use examples for style, label them as illustrative and vary them.
A real-life example
The ticket classifier from the zero-shot lesson still confuses two labels: a customer who says "cancel and return my money" for an undelivered order. The business rule is that an undelivered order is always delivery, even if the customer asks for money back. Definitions did not fix it, so the team adds targeted examples:
Classify the ticket. Return only the label.<example>Ticket: Order never arrived, please refund me.Label: delivery</example><example>Ticket: Food was cold and I want my money back.Label: refund</example><example>Ticket: Charged Rs 450 twice for one order.Label: billing</example>Ticket: Rider marked it delivered but I got nothing. Cancel and refund.Label:Before, this ticket came back as refund. After, it is delivery, and on 300 test tickets the delivery-versus-refund confusion drops from 31 errors to 6. The three examples add about 90 tokens per call, which the team judges worth it.
Follow-up questions to expect
- "How many examples should you use?" — Start with two to five and measure. Gains usually flatten quickly; beyond that, more examples mainly add cost.
- "Does the order of examples matter?" — It can: models lean towards the label of the last example. Shuffle or balance, and test with a different order.
- "Few-shot or fine-tuning?" — Few-shot for a handful of patterns you can change in minutes; fine-tuning when you have thousands of examples and need the behaviour cheaply at high volume.