Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
What you will build: PolicyPal, from prompt to production
Harbourline is a 2,000-person software company with offices in Bengaluru, Pune and Manchester. Its HR helpdesk receives about 1,100 tickets a month, and its IT helpdesk about 1,400. When the helpdesk leads went through a month of tickets, they found that roughly 60% of HR tickets and 35% of IT tickets were questions already answered somewhere in one of the company's 400 policy PDFs. People could not find the right PDF, or found the 2022 version, or could not tell whether a rule for India also applied in the UK.
So the ask is simple to say and hard to do well: an internal assistant that answers HR and IT policy questions correctly, shows where the answer came from, can check a person's own leave balance, can raise an IT ticket, and never leaks one employee's data to another. That assistant is PolicyPal, and you will build all of it in this course.
This first lesson shows you the finished product, so every later lesson has a place to land. You will see one real conversation, the numbers the system must hit, and the map of how it grows from a single prompt to a traced, evaluated and guarded service.
A conversation with the finished PolicyPal
Here is Priya, an engineer in Pune, talking to the version you will have at the end of Section 9.
Priya: I'm back from maternity leave on 3 March. Can I work from home three days a week for the first month? Also my new laptop won't connect to VPN.PolicyPal: Yes. The Return-to-Work policy allows up to four work-from-home days a week for the first eight weeks after parental leave, with your manager's approval [1]. You need to record the arrangement in the HR portal before your first day back [1]. For VPN on a new laptop, the device must be enrolled in device management before the VPN profile installs [2]. I can raise an IT ticket for this. Shall I raise it as "VPN not connecting on new laptop", normal priority? [1] Return-to-Work and Flexible Hours Policy (India), v3.2, s4.2 [2] Remote Access Standard, v5.0, s2.1Priya: Yes please.PolicyPal: Done. Ticket IT-48213 is raised. The IT team replies within one working day for normal priority tickets [2].A lot happened in those few seconds. A small fine-tuned router decided this message needed both policy retrieval and a ticket tool. A hybrid search over roughly 6,200 policy chunks found the two right sections, a reranker put them on top, and the answering model wrote a reply that cites only what it was given. Because raising a ticket changes something in the real world, the agent asked for confirmation first. Before the reply reached Priya, an output check confirmed that every sentence had a source and no other person's data was present. Every step wrote a span to a trace, so an engineer can later see why this answer was given.
Notice what PolicyPal did not do. It did not decide Priya's leave dates, it did not approve the work-from-home request, and it did not guess which VPN client she had. Knowing where the assistant stops is as much a part of the design as what it does.
The numbers the finished system has to hit
A demo is judged by whether it looks impressive. A production system is judged by numbers agreed before launch. These are PolicyPal's, and each one is earned in a specific section.
| Metric | Target | Where you earn it |
|---|---|---|
| Right policy section in the top 5 results (recall@5) | 0.90 or better | Section 3 |
| Answer sentences supported by a cited source | 97% or better | Sections 3 and 6 |
| Router accuracy on six routes | 95% or better | Section 5 |
| Sensitive HR topics sent to a human (recall) | 99% or better | Section 5 |
| Time to first visible word, median | under 1 second | Section 7 |
| Model cost per question | under $0.01 | Sections 1 and 7 |
| Red-team attacks that succeed | under 2% | Section 9 |
Two of these targets deserve a note now. The "sensitive topics" target is about harm, not convenience: a question about harassment or a medical condition should reach a trained HR partner, never an automated answer, so missing one in a hundred is already too many. And the cost target is not the hard part. At around 1,200 questions a working day, even a careless design costs a few hundred dollars a month. Correctness and safety are where the engineering effort goes.
Nine versions of one product
PolicyPal grows one section at a time. Each version exists because the previous one failed in a way you can measure.
| Section | PolicyPal gains | The failure it fixes |
|---|---|---|
| 1. Foundations | A chosen model behind a small client, cost logging | Guessing about models and budgets |
| 2. Prompts | A system prompt, a JSON answer contract, examples | Rambling answers the app cannot parse |
| 3. RAG | Search over 400 PDFs, citations | The model does not know Harbourline's rules |
| 4. Agents | Leave-balance and IT-ticket tools, a bounded loop | It can talk about a ticket but not raise one |
| 5. Fine-tuning | A small fine-tuned router | Routing with a large model is slow |
| 6. Evaluation | A 250-case eval suite and a careful judge | "It seems better" is not evidence |
| 7. LLMOps | A FastAPI service, caching, streaming, tracing | Nobody can see what it did in production |
| 8. Multimodal | Scans, screenshots and voice questions | Scanned policies are invisible to search; staff without laptops cannot type |
| 9. Safety | Red-team suite, PII controls, layered guardrails | One clever message leaks someone's data |
How the code is organised
You will write Python 3.11 or later throughout. The service uses FastAPI and Pydantic v2; retrieval uses sentence-transformers and rank_bm25; fine-tuning uses Hugging Face transformers and peft. The model provider sits behind one small class, so you can change providers by changing configuration.
policypal/ llm.py # the model client (Section 1) usage.py # token and cost accounting (Section 1) prompts/ # versioned system prompts and templates (Section 2) schemas.py # Pydantic models for answers and tool inputs (Section 2) ingest.py # PDF -> chunks -> embeddings (Section 3) retrieve.py # hybrid search and reranking (Section 3) tools.py # leave balance, IT tickets (Section 4) agent.py # the bounded agent loop (Section 4) router/ # fine-tuned routing model (Section 5) evals/ # eval sets, metrics, judge (Section 6) app.py # FastAPI service, streaming, tracing (Section 7) vision.py, speech.py # multimodal input (Section 8) guards.py # input, output and PII guardrails (Section 9)tests/Each file appears in the lesson that needs it. You will not write a framework first and use it later. That order is deliberate: every piece of code in this course exists because a real failure demanded it.
A final word on how to read the course. Each lesson ends with a short quiz whose explanations are part of the lesson, so read them even when you answer correctly. Where a lesson includes an exercise callout, do it with your own documents if you can. Retrieval, routing and evaluation all behave differently on real data, and the fastest way to learn them is to see your own numbers move.
Check your understanding
0 of 3 answered
1.Version 0 pasted twelve policies into one prompt and answered a Manchester employee with the India leave rule. What is the most useful way to treat Version 0?
2.Why is the target for sending sensitive HR topics to a human set at 99% recall, much stricter than the router's overall 95% accuracy?
3.PolicyPal asked Priya before raising the IT ticket, but did not ask before searching the policies. Why the difference?