Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you detect failures across multi-step AI pipelines?
What you need to know
Why multi-step pipelines hide failures
In a RAG or agent pipeline, errors do not stop the flow. Empty retrieval becomes a confident answer without facts. A failed tool call becomes "I could not find any orders" — which the user reads as fact. The final output looks normal, so status codes and error rates stay green.
Checks per step
| Step | Silent failure | Check |
|---|---|---|
| Retrieval | Zero chunks, or top score below threshold | retrieval.hit = false |
| Rerank | Nothing above the cutoff | rerank.kept = 0 |
| LLM call | finish_reason = length; empty output | Truncation and empty rates |
| Parsing | Invalid JSON or schema | Validation-failure rate |
| Tool call | Error, timeout, retry, unexpected empty result | Tool error and empty-result rates |
| Agent loop | Same tool called repeatedly; step limit hit | Loop and max-step counters |
| Output | Unsupported claim; missing citation | Groundedness or citation check |
The funnel dashboard
Plot the success rate of each step over time. A drop in final quality almost always appears first as one step's rate falling — for example retrieval hit rate going from 94% to 71%. That points you straight at the broken component.
Hard budgets
Give every request a maximum number of agent steps, a wall-clock limit and a token budget. A stuck agent should fail fast with a clear error, not loop 40 times and spend $5 on one question.
Replay
Store the inputs to each step (query, retrieved IDs, prompt version). When you fix something, re-run the stored failing inputs against the fix and compare outcomes before shipping.
A real-life example
A hospital's on-prem summariser has four steps: split the record into sections, extract medications (structured output), summarise each section, and combine. For two weeks, doctors report that some summaries "miss the medication list". The final judge score has dropped only from 4.3 to 4.1 — too small to alarm anyone.
The funnel dashboard tells the real story: section splitting 99.8%, medication extraction schema-valid 99.5%, but medication list non-empty fell from 97% to 82% on the day a new lab system went live. Traces show that the new system's PDFs label the section "Rx Summary" instead of "Medication Chart", so the splitter never sends that section to the extractor. A new rule for the label, replayed on 300 stored failing records, brings the rate back to 97%.
Follow-up questions to expect
- "How do you alert on so many step metrics?" — Alert on relative drops per step against the last week, and group them on one funnel dashboard.
- "How do you tell a retrieval failure from a model failure?" — Look at the retrieved context in the trace: if the right facts were not there, it is retrieval; if they were there and ignored, it is the prompt or model.
- "What should a tool return when it fails?" — An explicit error the model can see and report, never an empty result that looks like "no data".