Course Content
LangChain Mastery
7 sections · 109 lessons
How do you document LangChain applications?
What you need to know
LLM apps have behaviour that lives outside normal code: prompt wording, retrieval settings, model choice. A new engineer who only reads the Python will miss why things are the way they are — and may "tidy up" the one prompt line that was preventing a class of wrong answers.
What to write down
| Area | What to record |
|---|---|
| Prompts | The file, its version, and a comment for each rule: what failure it prevents |
| Data flow | Sources, loaders, chunk size and overlap, embedding model, index name, refresh schedule |
| Chain contracts | Exact input keys and output shape |
| Tools | What it does, when the model should use it, side effects, idempotency, permissions |
| Agents | Tools list, middleware (limits, approvals, fallbacks), stopping rules |
| Models | Primary and fallback models, typical tokens and cost per request, timeouts |
| Evaluation | Datasets, metrics, current scores, and what "good enough" means |
| Runbook | Common failures and what to do (provider outage, stale index, cost spike) |
Contracts in docstrings
1def build_rag_chain(retriever, llm) -> Runnable:2 """Answer HR policy questions with citations.34 Input: str (the employee's question)5 Output: {"question": str, "context": list[Document], "answer": str}6 Notes: answers only from context; says "I don't know" otherwise.7 Prompt: prompts/hr_answer_v7.txt8 """LCEL's dictionary plumbing is where new team members get stuck. A two-line contract saves them an hour of printing intermediate values.
Tools document themselves — for the model
A tool's docstring is sent to the model as its description. Write it for the model and for humans: what it does, when to use it, when not to. The same text appears in traces, so a clear docstring also helps debugging.
Diagrams and decisions
A short diagram of the request flow (API → agent → tools → databases) and a decision log ("switched to hybrid search on 3 March because ids were missed; recall@5 0.71 to 0.84") are worth more than pages of prose.
A real-life example
A travel-booking team's refund agent had a prompt rule: "Never promise a refund amount; say the airline will confirm." A new engineer removed it as "overly cautious" while shortening the prompt. Within a week the agent quoted refund amounts that airlines then refused, and support got angry customers.
The fix, beyond restoring the line, was documentation: every rule in prompts/refund_v4.txt now has a comment with the incident or ticket that caused it, the README describes each tool's side effects (cancel_booking is irreversible and requires human approval), and a docs/evals.md page lists the 80-case refund dataset and the minimum passing score. Prompt changes now require running that dataset in CI.
Follow-up questions to expect
- "Where should prompts live?" — In version control next to the code, or in a prompt registry with versions; either way, logged with every run.
- "How do you keep docs current?" — Keep them next to the code, make CI check that every tool has a docstring, and update the decision log in the same pull request as the change.
- "Do you document the model's behaviour?" — Yes, known limitations and failure cases, backed by eval results rather than opinions.