Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

What you will build: PolicyPal, from prompt to production


Harbourline is a 2,000-person software company with offices in Bengaluru, Pune and Manchester. Its HR helpdesk receives about 1,100 tickets a month, and its IT helpdesk about 1,400. When the helpdesk leads went through a month of tickets, they found that roughly 60% of HR tickets and 35% of IT tickets were questions already answered somewhere in one of the company's 400 policy PDFs. People could not find the right PDF, or found the 2022 version, or could not tell whether a rule for India also applied in the UK.

So the ask is simple to say and hard to do well: an internal assistant that answers HR and IT policy questions correctly, shows where the answer came from, can check a person's own leave balance, can raise an IT ticket, and never leaks one employee's data to another. That assistant is PolicyPal, and you will build all of it in this course.

This first lesson shows you the finished product, so every later lesson has a place to land. You will see one real conversation, the numbers the system must hit, and the map of how it grows from a single prompt to a traced, evaluated and guarded service.

Nine versions of one productPrompt behinda small clientJSONanswer contractRAG over 400policy PDFsTools and abounded agentFine-tuned0.5B router250-caseeval suiteServicewith tracingScans,screenshots and voiceLayered guardrailstopbottomVersion 0 — twelve policies pasted into one prompt — is the baseline every layer must beat.
Each version is built on the one below it, and each exists because the previous one failed on a number you can measure.

A conversation with the finished PolicyPal

Here is Priya, an engineer in Pune, talking to the version you will have at the end of Section 9.

Text
Priya:     I'm back from maternity leave on 3 March. Can I work from home           three days a week for the first month? Also my new laptop won't           connect to VPN.PolicyPal: Yes. The Return-to-Work policy allows up to four work-from-home           days a week for the first eight weeks after parental leave, with           your manager's approval [1]. You need to record the arrangement in           the HR portal before your first day back [1].           For VPN on a new laptop, the device must be enrolled in device           management before the VPN profile installs [2]. I can raise an           IT ticket for this. Shall I raise it as "VPN not connecting on new           laptop", normal priority?           [1] Return-to-Work and Flexible Hours Policy (India), v3.2, s4.2           [2] Remote Access Standard, v5.0, s2.1Priya:     Yes please.PolicyPal: Done. Ticket IT-48213 is raised. The IT team replies within           one working day for normal priority tickets [2].

A lot happened in those few seconds. A small fine-tuned router decided this message needed both policy retrieval and a ticket tool. A hybrid search over roughly 6,200 policy chunks found the two right sections, a reranker put them on top, and the answering model wrote a reply that cites only what it was given. Because raising a ticket changes something in the real world, the agent asked for confirmation first. Before the reply reached Priya, an output check confirmed that every sentence had a source and no other person's data was present. Every step wrote a span to a trace, so an engineer can later see why this answer was given.

Notice what PolicyPal did not do. It did not decide Priya's leave dates, it did not approve the work-from-home request, and it did not guess which VPN client she had. Knowing where the assistant stops is as much a part of the design as what it does.

The numbers the finished system has to hit

A demo is judged by whether it looks impressive. A production system is judged by numbers agreed before launch. These are PolicyPal's, and each one is earned in a specific section.

MetricTargetWhere you earn it
Right policy section in the top 5 results (recall@5)0.90 or betterSection 3
Answer sentences supported by a cited source97% or betterSections 3 and 6
Router accuracy on six routes95% or betterSection 5
Sensitive HR topics sent to a human (recall)99% or betterSection 5
Time to first visible word, medianunder 1 secondSection 7
Model cost per questionunder $0.01Sections 1 and 7
Red-team attacks that succeedunder 2%Section 9

Two of these targets deserve a note now. The "sensitive topics" target is about harm, not convenience: a question about harassment or a medical condition should reach a trained HR partner, never an automated answer, so missing one in a hundred is already too many. And the cost target is not the hard part. At around 1,200 questions a working day, even a careless design costs a few hundred dollars a month. Correctness and safety are where the engineering effort goes.

Nine versions of one product

PolicyPal grows one section at a time. Each version exists because the previous one failed in a way you can measure.

SectionPolicyPal gainsThe failure it fixes
1. FoundationsA chosen model behind a small client, cost loggingGuessing about models and budgets
2. PromptsA system prompt, a JSON answer contract, examplesRambling answers the app cannot parse
3. RAGSearch over 400 PDFs, citationsThe model does not know Harbourline's rules
4. AgentsLeave-balance and IT-ticket tools, a bounded loopIt can talk about a ticket but not raise one
5. Fine-tuningA small fine-tuned routerRouting with a large model is slow
6. EvaluationA 250-case eval suite and a careful judge"It seems better" is not evidence
7. LLMOpsA FastAPI service, caching, streaming, tracingNobody can see what it did in production
8. MultimodalScans, screenshots and voice questionsScanned policies are invisible to search; staff without laptops cannot type
9. SafetyRed-team suite, PII controls, layered guardrailsOne clever message leaks someone's data

How the code is organised

You will write Python 3.11 or later throughout. The service uses FastAPI and Pydantic v2; retrieval uses sentence-transformers and rank_bm25; fine-tuning uses Hugging Face transformers and peft. The model provider sits behind one small class, so you can change providers by changing configuration.

Text
policypal/  llm.py          # the model client (Section 1)  usage.py        # token and cost accounting (Section 1)  prompts/        # versioned system prompts and templates (Section 2)  schemas.py      # Pydantic models for answers and tool inputs (Section 2)  ingest.py       # PDF -> chunks -> embeddings (Section 3)  retrieve.py     # hybrid search and reranking (Section 3)  tools.py        # leave balance, IT tickets (Section 4)  agent.py        # the bounded agent loop (Section 4)  router/         # fine-tuned routing model (Section 5)  evals/          # eval sets, metrics, judge (Section 6)  app.py          # FastAPI service, streaming, tracing (Section 7)  vision.py, speech.py   # multimodal input (Section 8)  guards.py       # input, output and PII guardrails (Section 9)tests/

Each file appears in the lesson that needs it. You will not write a framework first and use it later. That order is deliberate: every piece of code in this course exists because a real failure demanded it.

A final word on how to read the course. Each lesson ends with a short quiz whose explanations are part of the lesson, so read them even when you answer correctly. Where a lesson includes an exercise callout, do it with your own documents if you can. Retrieval, routing and evaluation all behave differently on real data, and the fastest way to learn them is to see your own numbers move.

Check your understanding

0 of 3 answered

1.Version 0 pasted twelve policies into one prompt and answered a Manchester employee with the India leave rule. What is the most useful way to treat Version 0?

2.Why is the target for sending sensitive HR topics to a human set at 99% recall, much stricter than the router's overall 95% accuracy?

3.PolicyPal asked Priya before raising the IT ticket, but did not ask before searching the policies. Why the difference?