Course Content
Building AI Features in React
5 sections · 21 lessons
Few-Shot Examples and Constraints for Consistent Output
With a clear list of labels, the sentiment badge was right on easy reviews. "Best headphones I have owned" was positive every time. The trouble was the borderline ones. "Great sound, but I returned them because the left ear died after a week" came back positive in some runs and negative in others. "Fine for the price" wobbled between positive and mixed.
When the team looked closer, they found that two of their own engineers disagreed on the same reviews about 6 percent of the time. The model was not being random for no reason. The labels had no rule for the hard cases, so both the model and the humans were guessing.
This lesson fixes that in two steps: first write the rule, then show it with a few examples. Showing examples in a prompt is called few-shot prompting.
Decide the rule before writing examples
An example teaches a boundary only if you know where the boundary is. The ShopLens team wrote these definitions:
- Positive: the buyer is happy overall; any complaints are minor. A satisfied review with no strong feelings is positive.
- Mixed: clear praise and a clear problem, and neither dominates.
- Negative: the main experience was bad. A defect or a return makes a review at least mixed, and a product that stopped working is negative.
With these rules, the returned headphones are negative: the product stopped working. "Fine for the price" is positive: satisfied, no strong feelings. Human disagreement on the 200 test reviews dropped from 6 percent to 2 percent before any model was involved.
Examples teach the boundary
Now the prompt gets the rules and four short examples, each chosen to sit on a boundary.
1// server/prompts.ts (continued)2export const SENTIMENT_PROMPT_VERSION = "sentiment-v4";34export const SENTIMENT_SYSTEM = `You label product reviews for a sentiment badge.5Labels:6- positive: the buyer is happy overall; complaints are minor. Satisfied but unexcited is positive.7- mixed: clear praise and a clear problem, and neither dominates.8- negative: the main experience was bad. A product that stopped working is negative.9A defect or a return makes a review at least mixed.1011Reply with only JSON, one item per review, in input order:12{"items":[{"id":"r1","label":"positive"}]}1314Examples (invented products, not from the input):15"Clear sound, and the battery lasts my whole commute week. The case feels cheap." -> positive16"Loved the fit for a month, then the right bud stopped charging." -> negative17"Noise cancelling is superb but they pinch after an hour." -> mixed18"Fine for the price. Nothing special, nothing wrong." -> positive`;1920export function sentimentUserPrompt(reviews: Review[]): string {21 return `Label these reviews:\n${formatReviews(reviews)}`;22}Look at what each example teaches. The first shows that a minor complaint (the case) does not make a review mixed. The second shows that a failure outweighs a month of praise. The third is the model case for mixed. The fourth handles the "unexcited" case that was wobbling. There is no example of an obviously positive or obviously negative review, because the model never got those wrong.
What examples cost
Each example here is about 25 tokens, so four examples add about 100 tokens to a prompt of about 3,300 tokens for a page of 20 reviews. That is 3 percent more input. At $1 per million input tokens on the small model, it adds $0.0001 per call. For this feature, examples are almost free.
The arithmetic changes for long outputs. Examples of a full 60-word summary would be about 90 tokens each, and they cause a worse problem than cost: the model copies them. With two example summaries that both began "Most buyers praise...", nearly every generated summary began the same way. For long free text, describe the style in rules and skip whole examples, or use short fragments.
Three ways examples backfire
- Copying. As above, details and phrasing from examples appear in real outputs. Vary the examples and keep them clearly separate from real data.
- Label bias. If three of four examples are positive, the model leans positive on borderline cases. Keep the labels in your examples roughly balanced, and do not always put the same label last.
- Hidden disagreement. Examples cannot fix a rule you have not decided. If two engineers would label an example differently, the model will be inconsistent there too.
Constraints that code can enforce belong in code
Few-shot examples make output more consistent. They do not make it guaranteed. So the constraints that matter most are also enforced in code:
| Constraint | In the prompt | In code |
|---|---|---|
| Label is one of three values | Listed and shown in examples | SentimentLabel enum in Zod rejects anything else |
| One item per review | "one item per review, in input order" | The route keeps only ids that were requested |
| JSON only | Stated and shown | The parser strips code fences and rejects prose |
You may also have heard advice to lower the "temperature" setting to make output more repeatable. Some providers still accept it, and it can help. Several newer models no longer accept sampling settings at all, so do not build your consistency on it. Rules, examples and validation work on every model.
Check your understanding
0 of 3 answered
1.The badge wobbles between labels on reviews that mention a return. What should you do first?
2.You add two full example summaries to the summary prompt. Most new summaries now begin with the same four words as the examples. What happened?
3.Four examples of 25 tokens each are added to a 3,300-token badge prompt on the small model ($1 per million input tokens). Roughly what do they add per call?