Course Content
LLMOps & Deployment
6 sections · 40 lessons
What is the purpose of feature flags in AI deployments?
What you need to know
What flags are used for in LLM systems
- Canary ramp — a new prompt or model goes to 1%, then 5%, 25%, 50%, 100%, with a pause and a metrics check at each step.
- Per-tenant or per-region release — pilot customers first; regulated accounts last.
- A/B experiments — two arms, with metrics sliced by arm.
- Kill switch — turn off an agent's refund tool, or the whole AI feature, in seconds.
- Degradation switch — drop the reranker or switch to a smaller model under load.
Three rules that make flags work
- The flag holds a bundle, not a boolean.
canarymeans prompt v15 + model snapshot B + top-k 5, not five separate flags that can combine in untested ways. - Assignment is sticky. Hash the user ID so a user always gets the same arm. Random assignment per request gives one person two personalities in one conversation.
- Log the arm on every trace, so every metric can be split by variant.
1import hashlib23def bucket(user_id: str, flag: str) -> float:4 """Stable number in [0, 100) for this user and this flag."""5 h = hashlib.sha256(f"{flag}:{user_id}".encode()).hexdigest()6 return int(h[:8], 16) / 0x100000000 * 10078CONFIGS = {9 "stable": {"prompt": "support-v14", "model": "model-a-2026-03", "top_k": 5},10 "canary": {"prompt": "support-v15", "model": "model-b-2026-07", "top_k": 5},11}1213def pick_config(user_id: str, canary_percent: float) -> dict:14 arm = "canary" if bucket(user_id, "support-bot-v15") < canary_percent else "stable"15 return {"arm": arm, **CONFIGS[arm]}Including the flag name in the hash means each experiment gets an independent split, so the same 5% of users are not always the guinea pigs. Across 100,000 simulated users, this assigns about 4,900 to the canary at 5%. Managed tools — LaunchDarkly, Unleash, Statsig, or a config table with a short cache — do the same thing with a UI and an audit log.
Flags versus infrastructure canaries
A Kubernetes canary shifts traffic between two deployments of a service. A flag switches configuration inside one deployment. For prompt and model changes, flags are faster and finer: you can target a tenant, a language or a percentage without a new deploy.
A real-life example
A fintech support bot moves from model A to a newer model B with a revised prompt. The eval set passes. The team ramps with a flag: 1% for a day, then 5%.
At 5%, the dashboard sliced by arm shows the escalation-to-human rate for Hindi conversations rising from 11% to 14% on the canary, while English is flat. The on-call engineer sets the canary to 0% — it takes effect in 30 seconds, with no deploy. The investigation finds model B handles Hinglish (mixed Hindi and English) poorly with the new prompt. They add 30 Hinglish cases to the eval set, fix the prompt, and restart the ramp.
Follow-up questions to expect
- "How is a flag different from an A/B test?" — A flag is the mechanism; an A/B test is an experiment design (hypothesis, metric, sample size) that uses flags to assign arms.
- "What is shadow mode?" — The new variant runs on a copy of real requests but its answers are logged, not shown; it tests cost, latency and behaviour with no user risk.
- "What goes wrong with flags over time?" — Old flags pile up and create untested combinations; give every flag an owner and an expiry date.