Prompt Engineering Mastery

Course Content

Prompt Engineering Mastery

6 sections · 32 lessons

How can you design prompts that reduce bias and generate safe outputs?


Average urgency for the same 500 tickets3.83.83.23.75Open-ended promptListed factors onlyEnglishHinglishEach pair is the same complaint; only the language changes.
Counterfactual pairs turn a vague worry about bias into a number you can fix and keep testing.

What you need to know

Models learn from human text, so they absorb its patterns: which names sound "senior", which accents sound "urgent", which neighbourhoods sound "risky". A prompt cannot delete those patterns, but it can narrow what the model is allowed to use.

Prompt-level techniques

  1. Explicit criteria — "Rate urgency only on: money lost, service blocked, safety risk, time waiting." A closed list leaves less room for hidden factors.
  2. Name what to ignore — "Do not consider the customer's name, language, grammar or spelling."
  3. Blind the input — mask names and other identifiers before the call when the task does not need them. This is stronger than asking the model to ignore them.
  4. Require evidence — "Quote the part of the ticket that justifies the rating."
  5. Safe refusals — for harmful requests, a short refusal plus a helpful alternative, in wording you define.
  6. Narrow scope — narrow tasks drift less.

Testing: counterfactual pairs

Take real inputs and create copies where only one attribute changes — "Priya" to "Rahul", English to Hinglish, Mumbai to Patna. Run both and compare. Any systematic difference is bias you can measure, fix and re-test.

Beyond the prompt

  • Output filters and moderation for unsafe content.
  • Human review for decisions about people — hiring, credit, claims.
  • Monitoring of outcome rates by group over time.

A real-life example

A support team uses a model to score ticket urgency from 1 to 5. An audit of 500 ticket pairs — the same complaint written once in English and once in Hinglish — finds Hinglish tickets score 0.6 lower on average. Customers who write in Hinglish wait longer for help.

The original prompt:

Text
Rate the urgency of this support ticket from 1 to 5.

The revised prompt:

Text
Rate urgency from 1 to 5 using only these factors:- money lost or blocked (for example a failed UPI payment)- service unusable- safety risk- hours already waitingIgnore the language, spelling, grammar and tone of the ticket; tickets inHindi, English and Hinglish are equally important.Return {"urgency": n, "evidence": "<quote from the ticket>"}.

The gap on the same 500 pairs drops to 0.05, within run-to-run noise. The team adds the pairs to their eval set so any future prompt change is checked for the same bias.

Follow-up questions to expect

  • "Is telling the model 'don't be biased' enough?" — No. Explicit criteria, blinded inputs and counterfactual tests work; a general instruction mostly does not.
  • "How do you measure bias?" — Counterfactual pairs and outcome rates per group on a test set, compared before and after each change.
  • "What must never rely on a prompt alone?" — Decisions with legal or financial effect on people; they need human review and audit logs.