Course Content
RAG Systems
12 sections · 66 lessons
Why is configuring LangSmith crucial for complex applications?
What you need to know
Setup
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=faq-bot-prodWith these set, every LangChain or LangGraph call is traced. For your own functions, add the @traceable decorator:
1from langsmith import traceable23@traceable(run_type="retriever", name="search_faq")4def search_faq(question: str) -> list[dict]:5 ... # your own vector store call67@traceable(name="answer_question", metadata={"prompt_version": "v12"})8def answer_question(question: str, user_id: str) -> str:9 docs = search_faq(question)10 ...run_type="retriever" tells LangSmith to show the output as a list of documents. The metadata lets you filter traces later, for example all answers from prompt_version v12.
What it gives you
- The rendered prompt. The real text sent to the model, which is often not what you assumed: a template variable left empty, a chunk truncated, chunks in the wrong order.
- Per-step latency and cost, so you know whether the reranker or the model is slow.
- Datasets from traces. A failing production request can be added to a dataset in one click and replayed as a regression test.
- Experiments. Run two versions of the pipeline over the same dataset with evaluators and compare the scores side by side.
- Monitoring. Dashboards for error rate, latency, cost and feedback over time.
Things to decide
- Data sensitivity. Traces contain user questions and document text. Redact personal data, set retention, or self-host (LangSmith offers a self-hosted option; Langfuse and Phoenix are open source).
- Sampling. At high volume, trace a percentage of requests plus every error.
- Tags. Add
user_id,tenant_id, and version ids for prompt, index and model, so a support ticket maps to one exact trace.
A real-life example
An e-commerce product Q&A service gets a support ticket: "It told me the shoes are waterproof. They are not." Without tracing, the engineer can only guess. With LangSmith, they search by the product id and the user id from the ticket and open the trace in under a minute.
The retriever span shows three chunks: the official spec ("water-resistant upper"), a seller description ("perfect for rainy days"), and a customer review ("totally waterproof!"). The rendered prompt shows reviews were included with no label saying they are reviews. The fix is two lines: tag each chunk's source_type in the <doc> tag and tell the model to prefer official specs. The engineer adds this trace to the regression dataset, runs the evaluation, and the case now passes.
Follow-up questions to expect
- "Is LangSmith only for LangChain apps?" — No. The
@traceabledecorator and SDK work with any Python or TypeScript code, and it also accepts OpenTelemetry traces. - "What is the overhead?" — Traces are sent in the background, so request latency barely changes. The real costs are storage and privacy, which is why you sample and redact.
- "What would you alert on?" — Error rate, p95 latency, cost per query, rate of empty retrievals, and average faithfulness on the sampled judge.