Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is LLM observability, and why is it critical for production systems?


What you need to know

Why normal monitoring is not enough

Traditional APM watches errors, latency and throughput. An agent can pass all three and still:

  • quote a refund policy that does not exist,
  • call the right tool with the wrong customer ID,
  • loop 14 times and give up politely.

These are semantic failures. To catch them you need the content of each step and quality signals on top.

What to capture per run

  • A trace with a span for every model call, tool call and retrieval.
  • Model name and version, prompt version ID, sampling or reasoning-effort settings.
  • Tokens in and out, cached tokens, cost.
  • Tool arguments and results (with personal data redacted).
  • Guardrail triggers, approvals, and the final outcome.
  • User feedback and evaluation scores when available.

What to watch on a dashboard

MetricWhat it tells you
Task success rateIs the agent doing its job?
Steps per task (median and p95)Is it wandering or looping?
Budget-exhaustion rateHow often it hits the step or spend limit
Per-tool error rate and p95 latencyWhich dependency is hurting
Cost per successful taskThe real unit economics
Escalation and human-override rateWhere trust is breaking

Tooling

OpenTelemetry has GenAI semantic conventions, a shared set of attribute names for model and agent spans (still evolving). Using them keeps your data portable between backends such as Langfuse, LangSmith, Arize Phoenix, Braintrust or your own stack. Redact personal data on ingest, sample routine successful traces, and keep all failures; they become your next test cases.

A real-life example

A travel-booking agent showed 99.8% uptime and normal latency, yet complaints about wrong hotel dates rose over two weeks.

The traces showed it. For requests like "3 nights from Friday", the search_hotels span had check_out set one day early in about 6% of runs. All those runs used prompt version v14, deployed 16 days earlier, which had reworded the date instructions. Nothing had errored.

The team rolled back to v13, added 30 date-arithmetic cases to the golden set, and moved date calculation into a resolve_dates tool so the model no longer does it. They also added a monitor: "check_out minus check_in does not equal requested nights" now raises an alert.

Follow-up questions to expect

  • "What would you alert on?" — Sudden changes in success rate, escalation rate, schema failures, per-tool errors, and cost per task, compared with the previous week.
  • "How do you handle personal data in traces?" — Redact or hash it at ingest, restrict trace access, and set retention limits.
  • "Do you log everything?" — Log all failures and a sample of successes. At high volume, full content for every success is costly and rarely needed.