Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is tracing in LLM pipelines? What are spans?


One triage run as a span treeinvoke_agent — 11.8 schat, step 1 — 1.9 stool spansquery_metrics — 1.1 ssearch_logs —6.2 s, 2 retries
An average says triage is slow; the tree says one tool retried twice and ate half the run.

What you need to know

What an agent trace looks like

Text
invoke_agent  incident-triage              11.8 s   $0.21├─ chat  model=large  step 1                1.9 s   in 7,210  out 180├─ execute_tool  get_recent_deploys          0.4 s├─ execute_tool  query_metrics               1.1 s├─ execute_tool  search_logs                 6.2 s   retries=2  status=ok├─ chat  model=large  step 2                1.6 s   in 12,940 out 240└─ chat  model=large  step 3 (final)        0.6 s   in 13,400 out 310

Read it top-down. Total time is 11.8 s, and one tool, search_logs, took 6.2 s because it retried twice. Aggregate metrics would only say "triage is slow".

Good tracing practice

  • Propagate one trace ID, plus session and user IDs, through every call and sub-agent.
  • One span per model call and per tool call; record redacted inputs and outputs.
  • Record retries, timeouts and guardrail triggers as span events.
  • Put token counts and cost on model spans, so cost per step and per tool is free to compute.
  • Use the OpenTelemetry GenAI conventions (attributes such as the model name and input and output token counts) so you are not tied to one vendor.
  • Sample routine successes; keep every failure.

Why agents need it more than normal services

An ordinary API request might have 5 spans. An agent run with sub-agents and parallel branches can have 40 or more, and each run takes a different path. Without the tree, you cannot answer "why did this cost $1.40 when the median is $0.20?"

A real-life example

A procurement agent's cost per task doubled over a week. The average showed only the rise.

Grouping spans by tool showed that get_vendor_catalogue was now called 9 times per task instead of 2. Opening one trace showed the reason: a vendor's catalogue API had started returning page 1 of 10 by default, and the agent kept calling it for the next page, each time re-sending the growing context.

The fix: the tool now fetches all pages itself and returns only the matching items. Calls per task went back to 2 and cost returned to normal. The team added an alert on "calls per task per tool" for every tool.

Follow-up questions to expect

  • "Trace versus log?" — Logs are individual events. A trace links events and timings across one request's whole path, with parent-child structure.
  • "How do you trace across services or sub-agents?" — Pass the trace context in headers or message metadata, so the sub-agent's spans attach to the parent trace.
  • "What do you do with traces besides debugging?" — Build evaluation sets from failures, compute cost per step, and find slow or flaky tools.