Course Content
LLMOps & Deployment
6 sections · 40 lessons
What are the stages of an AI product lifecycle, from idea to deployment?
What you need to know
The stages and their exit gates
- Frame the problem — which user, which task, what a wrong answer looks like. Exit when you can describe failure precisely.
- Build the eval set first — 50–200 real inputs with expected behaviour. Teams that skip this cannot tell later whether they improved.
- Feasibility spike — simplest prompt on the strongest model. Exit if the best case is good enough to be worth engineering.
- Prototype — add retrieval, tools, structured output. Exit when it passes the eval set.
- Harden — guardrails, fallbacks, tracing, cost and latency budgets. Exit when p95 latency and cost per request are within budget.
- Shadow and canary — run on real traffic without users seeing it, then 1–5% live.
- General availability — full rollout behind a feature flag with a kill switch.
- Operate — sampled online evals, drift monitoring, incident runbooks, cost dashboards.
- Iterate or retire — new model versions re-enter at stage 5, not stage 1.
Why this order
- Eval before prototype, because a demo that looks good on five hand-picked inputs tells you nothing.
- Strongest model first, because it shows the ceiling. If even the best model fails, cheaper models will not save the project. Once it works, you try smaller models to cut cost.
- Hardening is its own stage, because most of the work between a demo and a product is here: timeouts, retries, fallbacks, guardrails, observability.
It is a loop, not a line
Production failures become new eval cases. Model upgrades, prompt changes and new data sources all go back through the eval gate and a canary. The eval set grows over the product's life and is the most valuable asset the team owns.
A real-life example
A hospital's discharge-summary project:
- Frame: doctors want a one-page summary. A wrong answer is any medicine or dose not in the source notes.
- Eval set: 120 past notes with summaries written by senior residents.
- Spike: a strong model, tested on de-identified notes only, hits 93% on the rubric — the ceiling is high enough.
- Prototype: they move to an open 8B model on-prem (the data rule) and add a section-by-section approach; it reaches 88%.
- Harden: an output check compares every medicine name in the summary to the source; unmatched names are flagged.
- Shadow: two weeks of summaries generated but shown only to reviewers; they find that the model ignores handwritten-note scans, so scans are routed to a human.
- Rollout: one ward first, then the whole hospital.
A year later a better open model arrives. It enters at the hardening stage, passes the same 120-case eval (now 400 cases), and ships through a canary.
Follow-up questions to expect
- "Where do most projects stall?" — Between prototype and production, usually because there is no eval set and nobody can prove it is good enough.
- "When would you retire an AI feature?" — When usage or quality metrics show it is not used or not trusted, or when its cost exceeds its value.
- "What goes into the eval set?" — Real inputs across main intents, known edge cases, and every production failure after launch.