LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is the purpose of feature flags in AI deployments?


Canary ramp for prompt v15 and model B1%5%25%50%100%01234Hindiescalations upnever reachedEscalations rose 11 to 14 percent for Hindi on the canary arm; the flag went to 0 in 30 seconds.
Because the arm was logged on every trace, a regression in one language was visible at 5 percent of traffic.

What you need to know

What flags are used for in LLM systems

  • Canary ramp — a new prompt or model goes to 1%, then 5%, 25%, 50%, 100%, with a pause and a metrics check at each step.
  • Per-tenant or per-region release — pilot customers first; regulated accounts last.
  • A/B experiments — two arms, with metrics sliced by arm.
  • Kill switch — turn off an agent's refund tool, or the whole AI feature, in seconds.
  • Degradation switch — drop the reranker or switch to a smaller model under load.

Three rules that make flags work

  1. The flag holds a bundle, not a boolean. canary means prompt v15 + model snapshot B + top-k 5, not five separate flags that can combine in untested ways.
  2. Assignment is sticky. Hash the user ID so a user always gets the same arm. Random assignment per request gives one person two personalities in one conversation.
  3. Log the arm on every trace, so every metric can be split by variant.
Python
import hashlibdef bucket(user_id: str, flag: str) -> float:    """Stable number in [0, 100) for this user and this flag."""    h = hashlib.sha256(f"{flag}:{user_id}".encode()).hexdigest()    return int(h[:8], 16) / 0x100000000 * 100CONFIGS = {    "stable": {"prompt": "support-v14", "model": "model-a-2026-03", "top_k": 5},    "canary": {"prompt": "support-v15", "model": "model-b-2026-07", "top_k": 5},}def pick_config(user_id: str, canary_percent: float) -> dict:    arm = "canary" if bucket(user_id, "support-bot-v15") < canary_percent else "stable"    return {"arm": arm, **CONFIGS[arm]}

Including the flag name in the hash means each experiment gets an independent split, so the same 5% of users are not always the guinea pigs. Across 100,000 simulated users, this assigns about 4,900 to the canary at 5%. Managed tools — LaunchDarkly, Unleash, Statsig, or a config table with a short cache — do the same thing with a UI and an audit log.

Flags versus infrastructure canaries

A Kubernetes canary shifts traffic between two deployments of a service. A flag switches configuration inside one deployment. For prompt and model changes, flags are faster and finer: you can target a tenant, a language or a percentage without a new deploy.

A real-life example

A fintech support bot moves from model A to a newer model B with a revised prompt. The eval set passes. The team ramps with a flag: 1% for a day, then 5%.

At 5%, the dashboard sliced by arm shows the escalation-to-human rate for Hindi conversations rising from 11% to 14% on the canary, while English is flat. The on-call engineer sets the canary to 0% — it takes effect in 30 seconds, with no deploy. The investigation finds model B handles Hinglish (mixed Hindi and English) poorly with the new prompt. They add 30 Hinglish cases to the eval set, fix the prompt, and restart the ramp.

Follow-up questions to expect

  • "How is a flag different from an A/B test?" — A flag is the mechanism; an A/B test is an experiment design (hypothesis, metric, sample size) that uses flags to assign arms.
  • "What is shadow mode?" — The new variant runs on a copy of real requests but its answers are logged, not shown; it tests cost, latency and behaviour with no user risk.
  • "What goes wrong with flags over time?" — Old flags pile up and create untested combinations; give every flag an owner and an expiry date.