Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is LLM observability, and why is it critical for production systems?
What you need to know
Why normal monitoring is not enough
Traditional APM watches errors, latency and throughput. An agent can pass all three and still:
- quote a refund policy that does not exist,
- call the right tool with the wrong customer ID,
- loop 14 times and give up politely.
These are semantic failures. To catch them you need the content of each step and quality signals on top.
What to capture per run
- A trace with a span for every model call, tool call and retrieval.
- Model name and version, prompt version ID, sampling or reasoning-effort settings.
- Tokens in and out, cached tokens, cost.
- Tool arguments and results (with personal data redacted).
- Guardrail triggers, approvals, and the final outcome.
- User feedback and evaluation scores when available.
What to watch on a dashboard
| Metric | What it tells you |
|---|---|
| Task success rate | Is the agent doing its job? |
| Steps per task (median and p95) | Is it wandering or looping? |
| Budget-exhaustion rate | How often it hits the step or spend limit |
| Per-tool error rate and p95 latency | Which dependency is hurting |
| Cost per successful task | The real unit economics |
| Escalation and human-override rate | Where trust is breaking |
Tooling
OpenTelemetry has GenAI semantic conventions, a shared set of attribute names for model and agent spans (still evolving). Using them keeps your data portable between backends such as Langfuse, LangSmith, Arize Phoenix, Braintrust or your own stack. Redact personal data on ingest, sample routine successful traces, and keep all failures; they become your next test cases.
A real-life example
A travel-booking agent showed 99.8% uptime and normal latency, yet complaints about wrong hotel dates rose over two weeks.
The traces showed it. For requests like "3 nights from Friday", the search_hotels span had check_out set one day early in about 6% of runs. All those runs used prompt version v14, deployed 16 days earlier, which had reworded the date instructions. Nothing had errored.
The team rolled back to v13, added 30 date-arithmetic cases to the golden set, and moved date calculation into a resolve_dates tool so the model no longer does it. They also added a monitor: "check_out minus check_in does not equal requested nights" now raises an alert.
Follow-up questions to expect
- "What would you alert on?" — Sudden changes in success rate, escalation rate, schema failures, per-tool errors, and cost per task, compared with the previous week.
- "How do you handle personal data in traces?" — Redact or hash it at ingest, restrict trace access, and set retention limits.
- "Do you log everything?" — Log all failures and a sample of successes. At high volume, full content for every success is costly and rarely needed.