Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Observability: traces, costs and quality signals


A message arrived in the PolicyPal feedback channel: "PolicyPal told me I can't claim my home broadband, but my colleague got a yes yesterday. Which is it?" Without observability, the team would have had a question, a vague time and two people's memories. They would have tried the question themselves, received some answer from today's system, and learned nothing.

With observability, the reply took four minutes. Each answer carries a trace id. The employee's trace showed the route, the retrieved chunks, the release, the model, the tokens and the grounding result. So did the colleague's. The colleague was in Manchester, and the UK home-working policy includes a broadband allowance; the India policy does not. PolicyPal had been right both times. The team replied with both citations, and the employee was satisfied.

Sometimes the trace shows PolicyPal was wrong, and then it shows exactly where. Either way, you cannot operate a system whose decisions you cannot see. This lesson makes PolicyPal's decisions visible.

One question as a span treepolicypal.ask4,410 msretrieve 182 msgenerate and checkdense 9 msrerank 149 msllm.answer 4,120 msgrounding 88 ms
Reading the span tree top down turns a complaint about an answer into the one stage that failed, in minutes.

What a trace records

A trace is the record of one request, made of spans: timed steps with attributes, nested like function calls. PolicyPal's trace for a policy question looks like this.

Text
policypal.ask                          4,410 ms  release=2026.03.18-2 route=policy_question  route                                   14 ms  label=policy_question p=0.97  retrieve                               182 ms    dense                                  9 ms  top_ids=[4410, 4411, 902, ...]    bm25                                  21 ms    rerank                               149 ms  top_score=6.3 chunk_ids=[4411, 4410, 3120, 902, 77]  llm.answer                           4,120 ms  model=claude-sonnet-5 in=3,040 out=231 usd=0.0084  grounding                               88 ms  sentences=3 unsupported=0

OpenTelemetry is the standard way to produce this. It is vendor-neutral: the same code can send traces to Jaeger, Grafana Tempo, Honeycomb, or LLM-focused tools such as Langfuse, by changing an exporter setting. Setup is a few lines at startup.

Python
# policypal/telemetry.pyfrom opentelemetry import tracefrom opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporterfrom opentelemetry.sdk.resources import Resourcefrom opentelemetry.sdk.trace import TracerProviderfrom opentelemetry.sdk.trace.export import BatchSpanProcessorprovider = TracerProvider(resource=Resource.create({"service.name": "policypal"}))provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))   # endpoint from envtrace.set_tracer_provider(provider)tracer = trace.get_tracer("policypal")

Then each stage of the pipeline wraps its work in a span and records what someone debugging it later would need.

Python
# inside Pipeline.answerwith tracer.start_as_current_span("policypal.ask") as root:    root.set_attribute("policypal.release", self.release)    with tracer.start_as_current_span("retrieve") as span:        chunks = self.retrieve(question, user.country)        span.set_attribute("policypal.chunk_ids", [c["id"] for c in chunks])        span.set_attribute("policypal.top_rerank", chunks[0]["rerank"])    with tracer.start_as_current_span("llm.answer") as span:        reply = self.llm.complete(self.system, messages, schema=ANSWER_SCHEMA, max_tokens=600)        span.set_attribute("gen_ai.request.model", self.llm.model)        span.set_attribute("gen_ai.usage.input_tokens", reply.input_tokens)        span.set_attribute("gen_ai.usage.output_tokens", reply.output_tokens)        span.set_attribute("policypal.cost_usd", record(self.llm.model, "answer", reply))

The gen_ai.* names follow OpenTelemetry's semantic conventions for generative AI, so tools that understand them can chart tokens and models without custom work. The policypal.* attributes are PolicyPal's own.

Notice what is not in the span: the question text and the answer text. Traces are widely readable by engineers and are kept for months. HR questions contain health, family and pay details. PolicyPal stores question and answer text in a separate, access-controlled store for 30 days, keyed by trace id, with names and IDs redacted. An engineer can look up the text for a specific trace when investigating, and that access is logged. Section 9 covers this in depth.

Cost and token dashboards

Because every model call records tokens and cost on its span, cost reporting becomes a query instead of a monthly surprise. PolicyPal's dashboard has four panels.

PanelWhy it exists
Cost per day, split by routeShows which feature spends the money; the agent path is 7% of traffic and about 18% of cost
Cost per question, by releaseCatches a prompt change that makes answers longer
Tokens per request, p50 and p99A rising p99 means some requests are growing, often an agent loop
Cache hit ratesAnswer cache and provider prompt cache, to confirm they still work

Two alerts sit on top: daily cost above twice its 7-day median, and any single request above 15,000 tokens. The second one caught a bug within an hour of a release: a change to history trimming stopped dropping old turns, and long conversations grew by about 3,000 tokens per turn.

Quality signals from production

Traces are also the raw material for quality monitoring. From them, a daily job computes the signals from Section 6, plus a few that come directly from the pipeline.

  • Re-ask rate and 24-hour ticket rate, the online quality metrics.
  • Status mix: the share of not_in_policy and needs_human answers.
  • Grounding drops: the share of answers where the NLI check removed a sentence or forced a regeneration.
  • Router confidence: the share of messages below the 0.6 confidence line.
  • Top rerank score: the average relevance score of the best retrieved chunk.

That last one is an early-warning signal for retrieval drift. When employees start asking about something the index does not contain, the best chunk they get is a weak match, and the top rerank score falls days before anyone complains.

A weekly review queue combines three sources: 50 random answers, every thumbs-down, and every answer that fell back to needs_human after a failed regeneration. Reviewers grade them with the Section 6 rubric, and every confirmed failure becomes an eval case.

Debugging a bad answer, step by step

  1. Find the trace — from the answer's trace id, shown in the UI's "report a problem" link.
  2. Read the spans top down — route and confidence, retrieved chunks and scores, the model's tokens and status, the grounding result.
  3. Decide which stage failed — wrong route, missing chunk, chunk present but misused, or a check that let it through.
  4. Replay with the same release — pin the release manifest from the trace and re-run the question, to confirm you can reproduce it.
  5. Add an eval case, then fix — the case fails before the fix and passes after, and the flip list shows nothing else broke.

Step 4 is only possible because of the release manifest from the first lesson. Without it, "replay" means running today's system, which may behave differently for reasons that have nothing to do with the bug.

Check your understanding

0 of 3 answered

1.Why does PolicyPal keep question and answer text out of trace attributes?

2.The average top rerank score drops sharply over three days with no code change. What does it most likely indicate?

3.When debugging a bad answer, why replay it with the release manifest from its trace instead of today's release?