LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

What is graceful degradation, and how is it applied in AI systems?


What you need to know

A degradation ladder

  1. Full quality — large model, full retrieval, reranker, all tools.
  2. Smaller model — same pipeline, cheaper and faster model.
  3. Shorter pipeline — no reranker, fewer chunks, shorter prompt, lower max_tokens.
  4. Cached answer — exact or close match to a common question.
  5. Retrieval only — show the most relevant documents and links, no generation.
  6. Honest unavailability — clear message, retry option, or "we will notify you".

For agents, degrade by lowering the maximum number of steps and switching off optional tools before refusing requests.

What makes it real

  • Evaluate each level in advance. Run the eval set at every level so you know, for example, that level 2 scores 86% against 92% at level 1.
  • Trigger automatically. Circuit-breaker state, queue depth, latency SLO burn or budget limits move the level; feature flags let on-call force a level.
  • Label the response. Add a field such as degraded_level: 3 so the UI can show a notice and downstream systems do not treat the answer as full quality.
  • Monitor time at each level. A system at level 2 for 30% of every evening has a capacity problem, even if uptime shows 100%.

Degrade by importance

Not all traffic is equal. Keep payments and safety-related questions on full quality longest; degrade "tell me about this product" first.

A real-life example

A state government chatbot for a scholarship scheme runs on a self-hosted cluster. On the final application day, traffic is 18× normal from 4 pm to midnight.

The ladder, prepared a month earlier:

  • Level 1 → 2 at queue depth over 200: the 70B model is replaced by an 8B model for general questions (eval: 88% vs 93%). Eligibility questions stay on 70B.
  • Level 3 at queue depth over 600: no reranker, top-3 chunks, answers capped at 250 tokens.
  • Level 4: the 50 most common questions ("last date?", "documents needed?") are answered from a pre-approved cache in all 12 languages; this covers 35% of traffic that day.
  • Level 5: for the rest, the bot shows the three most relevant official FAQ entries with links.

Nobody sees an error page. The dashboard shows 6 hours at level 2 or worse, which becomes the argument for pre-scaling next year.

Follow-up questions to expect

  • "How is this different from a fallback?" — A fallback reacts to a failure of one call; degradation is a planned, system-wide mode under stress. Fallbacks are often steps on the ladder.
  • "How do you avoid flapping between levels?" — Use hysteresis: move down at one threshold and back up only at a lower one, after a hold time.
  • "Should users be told?" — Yes, briefly; it sets expectations and reduces repeated questions.