Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 9: Observability and Debugging


Scenario: a LangChain app in production gives a wrong answer, and nobody can reconstruct why. How do you make every answer debuggable?

What you need to know

"The bot said something wrong" is not debuggable. "Here is the trace for request 7f3a: the query was rewritten badly, so retrieval returned the 2023 policy" is. The goal is that every answer can be traced back to its inputs.

Turn it on

Bash
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=support-bot-prod

Every chain, retriever, tool and model call is then traced. For RAG, that already answers the key question: was the right chunk retrieved?

What to add beyond the default

  1. Correlate — pass metadata and tags on each call so a support ticket maps to a trace in one search.
  2. Capture feedback — log thumbs up and down against the run id through the LangSmith feedback API.
  3. Promote failures — bad traces become examples in an eval dataset, which runs in CI on every prompt or model change.
  4. Sample and mask — trace all errors and low-rated runs, sample the rest, and mask PII before data leaves your network.
Python
result = chain.invoke(    {"question": q},    config={"metadata": {"user_id": user.id, "tenant": user.tenant, "request_id": req_id},            "tags": ["prod", "prompt-v3"], "run_name": "support_answer"},)

The same request_id goes into your application logs, so a complaint, a log line and a trace link together.

What a trace tells you

What you see in the traceDiagnosis
Right chunk not in the retriever outputRetrieval problem: chunking, hybrid search, filters
Right chunk retrieved but the answer contradicts itGeneration problem: prompt, conflicting chunks, faithfulness
Tool returned an error or empty resultTool or data problem
Correct answer, but the user rated it downTone, length or expectation problem

When LangSmith isn't allowed

Some companies cannot send prompts to a third-party service. Options: self-host LangSmith, or send OpenTelemetry traces to your existing backend; LangChain and LangSmith support OpenTelemetry export, and open-source tools such as Langfuse accept LangChain traces. The non-negotiable part is capturing inputs and outputs per span. Without that, you are guessing.

A real-life example

Scenario (illustrative numbers). A mutual-fund platform's assistant tells a customer the exit load on a fund is 0%; it is 1% within a year. Compliance asks why. The team has only application logs with the question and answer.

After enabling tracing with request ids, the next similar complaint takes 10 minutes to diagnose: the retriever returned a factsheet for a different share class of the same fund, because the class name was in a table the chunker had flattened. The team fixes the chunking, adds 25 share-class questions to the eval dataset from traced failures, and sets tracing to capture 100% of thumbs-down runs and 10% of the rest, with PAN and account numbers masked before export.

Follow-up questions to expect

  • "Doesn't tracing add latency?" — Traces are sent in the background; the overhead per request is small. Sampling reduces volume and cost further.
  • "How long do you keep traces?" — Match your data-retention policy; keep full traces for a short period and aggregated metrics longer, and keep regulated audit data as required.
  • "How do traces feed evaluation?" — Failing and low-rated traces become labelled dataset examples, so the eval set grows from real failures.