Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Examples and reasoning: few-shot and step-by-step, used carefully


With the version 1 prompt and the JSON contract in place, PolicyPal passed most of its smoke set. Two failures remained, and both were stubborn. The first was about judgement. When a policy covered a related topic but not the exact question, for example "Can I carry forward unused sick leave?" when the policy only discussed carrying forward earned leave, the model chose answered in some runs and not_in_policy in others. On 60 such partial-coverage questions, it was consistent only 89% of the time.

The second failure was arithmetic. "I joined on 15 September. Earned leave is 18 days a year and accrues monthly. How many days will I have on 31 December?" The model got questions like this wrong 30% of the time, usually by one month.

Rewording the rules did not help either problem. This lesson covers the two tools people reach for next, examples and step-by-step reasoning, how to use each on exactly these failures, and the third option that beats both for arithmetic.

Leave accrual: who should do the arithmetic?Model calculates• 70% right with a plain prompt• 88% with a reasoning field• 180 more output tokens each• About 2.5 s slowerModel extracts, code computes• 97.5% right: 39 of 40• Reading dates is the only model step• Errors move to extraction, easy to see• Completed months counted exactly
Let the model do the language work of reading inputs and let code do the counting it never gets wrong.

Few-shot: showing beats telling at a boundary

Rules describe a boundary in words. Examples show the model where the boundary actually is. For a judgement call like "is this question covered?", three short examples placed near the boundary do more than a paragraph of explanation.

Text
<example>Sources: [1] "Unused earned leave, up to 30 days, is carried forward to the next year."Question: Can I carry forward sick leave?Reply: {"status": "not_in_policy", "answer": "The policy covers carrying forwardearned leave [1], but it does not say anything about sick leave. Please ask yourHR partner.", "citations": [1], "follow_up": ""}</example><example>Sources: [1] "Employees may claim one internet bill per month, up to ₹1,500."Question: Can I claim my broadband bill?Reply: {"status": "answered", "answer": "Yes. You can claim one internet bill amonth, up to ₹1,500 [1].", "citations": [1], "follow_up": ""}</example>

Good examples follow a few rules. They sit on the boundary, not in the easy middle: the first example is a near miss, the second a true match. They use invented policies, not real ones, so the model cannot copy a real fact from an example into the wrong answer. They are short, because every example is paid for on every call. And they are balanced: if all three examples say not_in_policy, the model learns that not_in_policy is the popular answer.

Three static examples added about 450 tokens per call, which is about $0.0009 per question at mid-tier prices. Consistency on the 60 partial-coverage questions went from 89% to 94%.

Dynamic examples: the right three for this question

A fixed set of three examples cannot cover every kind of boundary. The next step is a bank of about 40 short, reviewed examples and a function that picks the three most similar to the current question.

Python
# policypal/examples.pyimport jsonimport numpy as npfrom sentence_transformers import SentenceTransformerembedder = SentenceTransformer("BAAI/bge-small-en-v1.5")BANK = [json.loads(line) for line in open("policypal/prompts/example_bank.jsonl")]BANK_VECS = embedder.encode([e["question"] for e in BANK], normalize_embeddings=True)def pick_examples(question: str, k: int = 3) -> list[dict]:    q = embedder.encode([question], normalize_embeddings=True)[0]    scores = BANK_VECS @ q                     # cosine similarity, since vectors are normalised    chosen = [BANK[i] for i in np.argsort(-scores)[: k * 2]]    # keep at least one "answered" and one not-answered example when possible    answered = [e for e in chosen if e["status"] == "answered"][:1]    others = [e for e in chosen if e["status"] != "answered"][:1]    rest = [e for e in chosen if e not in answered + others]    return (answered + others + rest)[:k]

The function takes the six nearest examples and then makes sure at least one "answered" and one "not answered" example survive, so the picked set is never all one label. The embedding model is small and runs in a few milliseconds on a CPU. You will use the same model for retrieval in Section 3, so this adds no new dependency.

With dynamic examples, consistency on the partial-coverage set rose to 97%. The extra cost was the same 450 tokens, because the number of examples did not change. Only their relevance did.

Step-by-step reasoning, and what it costs

For the leave calculation, the classic advice is to ask the model to "think step by step" before answering. It works because the model produces its answer one token at a time: if the working is written first, the final number is conditioned on that working instead of being guessed in one step.

There are two ways to get this today. You can add a working field to the schema, placed before the answer, so the model writes its reasoning in visible text. Or you can use a model's built-in reasoning mode, sometimes called extended thinking, which reasons before replying and is usually controlled by an effort setting (for example output_config={"effort": "low"} on the Anthropic API). Both work, and both cost the same thing: reasoning tokens are output tokens.

For PolicyPal's calculation questions, a working field raised accuracy from 70% to 88%. It also added about 180 output tokens per question, roughly $0.0018 and 2.5 seconds. That is acceptable for the 3% of questions that involve a calculation. It would be a poor trade on the 97% that are simple lookups, where reasoning adds latency and, in the pilot, occasionally "reasoned" its way past the source into general knowledge.

Better still: let code do the arithmetic

88% is better than 70%, but it still means one calculation in eight is wrong, and leave balances are exactly the kind of number people act on. The strongest fix is to stop asking the model to calculate at all. Ask it to extract the inputs, which it does very reliably, and let Python do the maths, which it does perfectly.

Python
# policypal/accrual.pyfrom datetime import date, timedeltafrom typing import Literalfrom pydantic import BaseModel, ConfigDictclass AccrualInputs(BaseModel):    model_config = ConfigDict(extra="forbid")    annual_days: float                 # from the policy text    join_date: str                     # "YYYY-MM-DD", from the question    as_of: str                         # "YYYY-MM-DD", from the question or today    accrual: Literal["monthly", "yearly"]def completed_months(start: date, end: date) -> int:    after = end + timedelta(days=1)    months = (after.year - start.year) * 12 + (after.month - start.month)    return months - (1 if after.day < start.day else 0)def accrued_days(inp: AccrualInputs) -> float:    start, end = date.fromisoformat(inp.join_date), date.fromisoformat(inp.as_of)    if inp.accrual == "yearly":        return inp.annual_days if end >= start else 0.0    return round(completed_months(start, end) * inp.annual_days / 12, 1)

For "joined 15 September, 18 days a year, monthly, as of 31 December", completed_months returns 3 and accrued_days returns 4.5. The model extracts AccrualInputs with the same structured-output call you already have, then writes the answer around the computed number, citing the policy for the rule. On the 40 calculation questions in the pilot, this path was right on 39. The one miss was an extraction error, a date written as "15/9" that the model read as 9 May until the prompt specified day-first dates for India. In Section 4 this function becomes a tool the agent can call.

When examples and reasoning make things worse

Neither tool is free, and both can hurt.

  • Copied examples. An early example used a real ₹1,500 internet limit. For a week, the model quoted ₹1,500 for a mobile bill question, copying the example instead of the source. Invented values fixed it.
  • Skewed labels. A bank with twice as many not_in_policy examples made the model more cautious everywhere. Balance the bank.
  • Reasoning on easy questions. It adds seconds and cost, and gives the model more room to drift from the source.
  • Leaked reasoning. Visible working can repeat internal rules or other context the user should not see. Keep it out of the displayed answer.

Measure each change on the same question set before and after. If you cannot show a number moved, remove the change.

Check your understanding

0 of 3 answered

1.An example in PolicyPal's prompt uses a real ₹1,500 internet allowance, and the model starts quoting ₹1,500 for mobile bills. What is the best fix?

2.Calculation questions are 3% of PolicyPal's traffic. Why not add a working field to every answer, just in case?

3.For "how many leave days will I have on 31 December?", why does extracting inputs and computing in Python beat asking the model to reason carefully?