Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is LLMOps, and how is it different from traditional MLOps practices?
What you need to know
What stays the same
MLOps gave us the habits that still apply: everything is versioned, every change goes through a pipeline with tests, production is monitored, and every release can be rolled back. An LLMOps engineer still needs all of these.
What changes
| Traditional MLOps | LLMOps | |
|---|---|---|
| Main artefact | Trained weights from your own data | Prompt + model snapshot ID + retrieval index + tool schemas |
| Training | Core of the job | Rare; sometimes a LoRA fine-tune |
| Measuring quality | Accuracy, F1, AUC on labelled data | Eval sets, rubric scoring, LLM-as-judge, human review |
| Cost driver | Training compute; cheap inference | Inference, billed per input and output token |
| Where drift comes from | Input data changes | Input changes, stale retrieval corpus, provider model updates |
| Typical failure | Wrong prediction | Fluent, confident, wrong answer with a 200 OK |
Three differences change daily engineering the most:
- The unit of release is a bundle. A "release" is a prompt version plus a model snapshot plus an index version plus sampling settings. Change any one and behaviour changes, so all four are versioned and logged together.
- Outputs are open-ended. There is no single correct string. You need a golden set of real cases and a scoring method (rules, a judge model, or people).
- Cost and latency vary per request. A long document in the prompt can make one request 20 times more expensive than another. Cost per request becomes a production metric, like p95 latency.
What this means in practice
An LLMOps setup has eval gates in CI, immutable prompt versions, pinned model IDs, token and cost data on every trace, guardrails in the request path, and gradual rollouts (shadow, then canary) instead of big-bang releases.
A real-life example
A fintech company already runs a fraud model with classic MLOps: a feature store, weekly retraining, AUC tracked on a dashboard. It now launches a customer-support chatbot on a hosted LLM.
The team's first week shows the difference. Nothing is retrained. Instead, a prompt edit that made answers "friendlier" also made the bot promise refunds it cannot give — the logs showed only HTTP 200s. A month later the provider moved the model alias to a new snapshot and the Hindi answers became longer and more expensive, with no deploy on their side.
Their fixes are pure LLMOps: a 300-case eval set run on every prompt change, the model pinned to a dated snapshot, an output guardrail that blocks refund promises, and a dashboard of cost per conversation by language.
Follow-up questions to expect
- "Do MLOps skills still matter?" — Yes. CI/CD, versioning, monitoring and rollback are the base. LLMOps adds prompt and eval management, token economics and guardrails on top.
- "When does LLMOps include training?" — When you fine-tune (usually LoRA) or distil a smaller model. Then the classic training pipeline comes back for that part.
- "What is the hardest part?" — Evaluation. Without a trusted eval set you cannot tell whether any change made things better or worse.