LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you detect failures across multi-step AI pipelines?


Per-step success rates in the summariser (percent)99.899.582012section splitschema validmedsfound, was 97The end judge score only slipped from 4.3 to 4.1. A renamed Rx Summary section was never extracted.
The end score barely moved while one step fell 15 points — the funnel names the broken step, the average hides it.

What you need to know

Why multi-step pipelines hide failures

In a RAG or agent pipeline, errors do not stop the flow. Empty retrieval becomes a confident answer without facts. A failed tool call becomes "I could not find any orders" — which the user reads as fact. The final output looks normal, so status codes and error rates stay green.

Checks per step

StepSilent failureCheck
RetrievalZero chunks, or top score below thresholdretrieval.hit = false
RerankNothing above the cutoffrerank.kept = 0
LLM callfinish_reason = length; empty outputTruncation and empty rates
ParsingInvalid JSON or schemaValidation-failure rate
Tool callError, timeout, retry, unexpected empty resultTool error and empty-result rates
Agent loopSame tool called repeatedly; step limit hitLoop and max-step counters
OutputUnsupported claim; missing citationGroundedness or citation check

The funnel dashboard

Plot the success rate of each step over time. A drop in final quality almost always appears first as one step's rate falling — for example retrieval hit rate going from 94% to 71%. That points you straight at the broken component.

Hard budgets

Give every request a maximum number of agent steps, a wall-clock limit and a token budget. A stuck agent should fail fast with a clear error, not loop 40 times and spend $5 on one question.

Replay

Store the inputs to each step (query, retrieved IDs, prompt version). When you fix something, re-run the stored failing inputs against the fix and compare outcomes before shipping.

A real-life example

A hospital's on-prem summariser has four steps: split the record into sections, extract medications (structured output), summarise each section, and combine. For two weeks, doctors report that some summaries "miss the medication list". The final judge score has dropped only from 4.3 to 4.1 — too small to alarm anyone.

The funnel dashboard tells the real story: section splitting 99.8%, medication extraction schema-valid 99.5%, but medication list non-empty fell from 97% to 82% on the day a new lab system went live. Traces show that the new system's PDFs label the section "Rx Summary" instead of "Medication Chart", so the splitter never sends that section to the extractor. A new rule for the label, replayed on 300 stored failing records, brings the rate back to 97%.

Follow-up questions to expect

  • "How do you alert on so many step metrics?" — Alert on relative drops per step against the last week, and group them on one funnel dashboard.
  • "How do you tell a retrieval failure from a model failure?" — Look at the retrieved context in the trace: if the right facts were not there, it is retrieval; if they were there and ignored, it is the prompt or model.
  • "What should a tool return when it fails?" — An explicit error the model can see and report, never an empty result that looks like "no data".