Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is tracing in LLM pipelines? What are spans?
What you need to know
What an agent trace looks like
invoke_agent incident-triage 11.8 s $0.21├─ chat model=large step 1 1.9 s in 7,210 out 180├─ execute_tool get_recent_deploys 0.4 s├─ execute_tool query_metrics 1.1 s├─ execute_tool search_logs 6.2 s retries=2 status=ok├─ chat model=large step 2 1.6 s in 12,940 out 240└─ chat model=large step 3 (final) 0.6 s in 13,400 out 310Read it top-down. Total time is 11.8 s, and one tool, search_logs, took 6.2 s because it retried twice. Aggregate metrics would only say "triage is slow".
Good tracing practice
- Propagate one trace ID, plus session and user IDs, through every call and sub-agent.
- One span per model call and per tool call; record redacted inputs and outputs.
- Record retries, timeouts and guardrail triggers as span events.
- Put token counts and cost on model spans, so cost per step and per tool is free to compute.
- Use the OpenTelemetry GenAI conventions (attributes such as the model name and input and output token counts) so you are not tied to one vendor.
- Sample routine successes; keep every failure.
Why agents need it more than normal services
An ordinary API request might have 5 spans. An agent run with sub-agents and parallel branches can have 40 or more, and each run takes a different path. Without the tree, you cannot answer "why did this cost $1.40 when the median is $0.20?"
A real-life example
A procurement agent's cost per task doubled over a week. The average showed only the rise.
Grouping spans by tool showed that get_vendor_catalogue was now called 9 times per task instead of 2. Opening one trace showed the reason: a vendor's catalogue API had started returning page 1 of 10 by default, and the agent kept calling it for the next page, each time re-sending the growing context.
The fix: the tool now fetches all pages itself and returns only the matching items. Calls per task went back to 2 and cost returned to normal. The team added an alert on "calls per task per tool" for every tool.
Follow-up questions to expect
- "Trace versus log?" — Logs are individual events. A trace links events and timings across one request's whole path, with parent-child structure.
- "How do you trace across services or sub-agents?" — Pass the trace context in headers or message metadata, so the sub-agent's spans attach to the parent trace.
- "What do you do with traces besides debugging?" — Build evaluation sets from failures, compute cost per step, and find slow or flaky tools.