LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How is CI/CD adapted for AI applications compared to standard software systems?


What a prompt change must pass before users see itChange: code,prompt, model or indexDeterministictests:parsers, toolsEval gate,scored perlanguageCost andlatencybudget checkCanary at5%, then rampEnglish rose 91 to 94, Tamil fell 86 to 79 — the per-language gate blocked it.
An average can improve while one slice breaks, so the gate compares every slice against a tolerance near the eval set's noise.

What you need to know

What triggers the pipeline

In normal software only code changes trigger CI. In an LLM app, four kinds of change alter behaviour and must all go through CI: code, prompt text, model snapshot ID, and retrieval config or index. Prompts and eval datasets live in the repo and go through pull-request review like code.

The pipeline

  1. Deterministic tests — unit and integration tests for parsers, tool wrappers, routing and schema validation. Fast and exact.
  2. Eval gate — run the golden set (100–500 labelled real cases) against the candidate and the current production baseline; score with rules, an LLM judge, or both.
  3. Budget checks — block if cost per request or p95 latency rises beyond a limit.
  4. Release — deploy behind a feature flag; shadow traffic or a 1–5% canary first, then ramp.
  5. Watch and roll back — compare live metrics by variant; flip the flag back if they worsen.

Why tolerance bands, and how wide

Outputs vary between runs even at temperature 0, and the eval set is a sample. With 300 cases and a 90% pass rate, the standard error is about 1.7 percentage points (sqrt(0.9 × 0.1 / 300)). So a drop from 91% to 90% is noise; a drop from 91% to 86% is real. Set the tolerance near this noise level, and run the suite two or three times if results jump around.

Gate on slices, not only the average

An average can rise while one important group falls. Score per language, per intent or per customer tier, and gate on each slice.

Python
import sysTOLERANCE = 0.02          # allow a 2-point dip before blockingMAX_COST_RISE = 0.10      # and at most 10% more cost per requestbase = {"pass_rate": {"en": 0.91, "hi": 0.88, "ta": 0.86}, "cost": 0.0041}cand = {"pass_rate": {"en": 0.94, "hi": 0.89, "ta": 0.79}, "cost": 0.0043}failures = []for lang, old in base["pass_rate"].items():    new = cand["pass_rate"][lang]    if new < old - TOLERANCE:        failures.append(f"{lang}: {old:.0%} -> {new:.0%}")if cand["cost"] > base["cost"] * (1 + MAX_COST_RISE):    failures.append("cost per request rose more than 10%")if failures:    print("BLOCKED:", "; ".join(failures))    sys.exit(1)print("passed")

This prints BLOCKED: ta: 86% -> 79% and exits with code 1, which fails the CI job. In a real pipeline the two dictionaries come from the eval runner's output files. Tools such as promptfoo, DeepEval or LangSmith evaluations produce these scores; the gate itself stays this simple.

A real-life example

A state government runs a multilingual chatbot for its pension scheme in English, Hindi and Tamil. A developer rewrites the system prompt to make answers shorter. On the English eval cases, the pass rate rises from 91% to 94%, and in a single-average pipeline the change would ship.

The per-language gate blocks it: Tamil falls from 86% to 79%. Reading the failed cases shows why — the shorter prompt dropped the line "answer in the user's language", and the model now replies in English to many Tamil questions. The developer restores the line, the gate passes, and the change goes to a 5% canary for three days before full rollout.

Follow-up questions to expect

  • "How big should the golden set be?" — Start with 100–200 real cases covering main intents and past failures; grow it with every production incident.
  • "How do you handle the cost of running evals in CI?" — Run a small fast subset on every commit and the full suite on merge to main or on prompt/model changes.
  • "Can you trust an LLM judge in CI?" — Only after checking it agrees with human labels on a sample, and with the judge model and prompt pinned.