LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you document LangChain applications?


What you need to know

LLM apps have behaviour that lives outside normal code: prompt wording, retrieval settings, model choice. A new engineer who only reads the Python will miss why things are the way they are — and may "tidy up" the one prompt line that was preventing a class of wrong answers.

What to write down

AreaWhat to record
PromptsThe file, its version, and a comment for each rule: what failure it prevents
Data flowSources, loaders, chunk size and overlap, embedding model, index name, refresh schedule
Chain contractsExact input keys and output shape
ToolsWhat it does, when the model should use it, side effects, idempotency, permissions
AgentsTools list, middleware (limits, approvals, fallbacks), stopping rules
ModelsPrimary and fallback models, typical tokens and cost per request, timeouts
EvaluationDatasets, metrics, current scores, and what "good enough" means
RunbookCommon failures and what to do (provider outage, stale index, cost spike)

Contracts in docstrings

Python
def build_rag_chain(retriever, llm) -> Runnable:    """Answer HR policy questions with citations.    Input:  str (the employee's question)    Output: {"question": str, "context": list[Document], "answer": str}    Notes:  answers only from context; says "I don't know" otherwise.            Prompt: prompts/hr_answer_v7.txt    """

LCEL's dictionary plumbing is where new team members get stuck. A two-line contract saves them an hour of printing intermediate values.

Tools document themselves — for the model

A tool's docstring is sent to the model as its description. Write it for the model and for humans: what it does, when to use it, when not to. The same text appears in traces, so a clear docstring also helps debugging.

Diagrams and decisions

A short diagram of the request flow (API → agent → tools → databases) and a decision log ("switched to hybrid search on 3 March because ids were missed; recall@5 0.71 to 0.84") are worth more than pages of prose.

A real-life example

A travel-booking team's refund agent had a prompt rule: "Never promise a refund amount; say the airline will confirm." A new engineer removed it as "overly cautious" while shortening the prompt. Within a week the agent quoted refund amounts that airlines then refused, and support got angry customers.

The fix, beyond restoring the line, was documentation: every rule in prompts/refund_v4.txt now has a comment with the incident or ticket that caused it, the README describes each tool's side effects (cancel_booking is irreversible and requires human approval), and a docs/evals.md page lists the 80-case refund dataset and the minimum passing score. Prompt changes now require running that dataset in CI.

Follow-up questions to expect

  • "Where should prompts live?" — In version control next to the code, or in a prompt registry with versions; either way, logged with every run.
  • "How do you keep docs current?" — Keep them next to the code, make CI check that every tool has a docstring, and update the decision log in the same pull request as the change.
  • "Do you document the model's behaviour?" — Yes, known limitations and failure cases, backed by eval results rather than opinions.