Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement unit tests for LangChain chains?
What you need to know
The testing layers
| Layer | Model | Checks | When |
|---|---|---|---|
| Unit | Fake | Prompt rendering, parsing, routing, filters, error handling | Every commit, seconds |
| Integration | Real, few calls | Provider wiring, structured output with the real model | Nightly or before release |
| Evaluation | Real, dataset | Quality: correctness, groundedness, tool choice | On prompt/model/retrieval changes |
A unit test with fakes
1from langchain_core.documents import Document2from langchain_core.language_models.fake_chat_models import FakeListChatModel3from langchain_core.runnables import RunnableLambda45def test_rag_puts_retrieved_text_in_prompt():6 seen = {}7 def spy(prompt_value): # records the prompt sent to the model8 seen["prompt"] = prompt_value.to_string()9 return prompt_value1011 llm = FakeListChatModel(responses=["You get 30 days [refunds.md]"])12 retriever = RunnableLambda(lambda q: [13 Document(page_content="Refunds within 30 days.", metadata={"source": "refunds.md"})])1415 chain = build_rag_chain(retriever, RunnableLambda(spy) | llm)16 out = chain.invoke("What is the refund window?")1718 assert "Refunds within 30 days." in seen["prompt"]19 assert "[refunds.md]" in out["answer"]20 assert out["context"][0].metadata["source"] == "refunds.md"- The fake model returns its responses in order, with no network.
- The spy step checks what the model would have received — the most important thing in a RAG test.
- Assertions are about structure (source present, context passed), not wording.
What else to unit test
- Parsers and validation — feed malformed text or objects; assert a clean error or repair.
- Tools — call the tool function directly with valid and invalid arguments (
tool.invoke({...})); mock the external API. - Agents — use a fake chat model that returns an
AIMessagewithtool_calls, then assert the right tool ran with the right arguments and that its result reached the next model turn. - Security — the retriever for tenant A never returns tenant B's documents.
- Error paths — make a fake raise
TimeoutErrorand assert the fallback message.
Quality in CI
With pip install "langsmith[pytest]", tests marked @pytest.mark.langsmith log inputs, outputs and feedback to LangSmith, so an evaluation suite can run as ordinary pytest with results tracked over time.
A real-life example
A finance assistant has a get_stock_price tool and an answer chain. Its tests used to call the real model and the real stock API: 4 minutes per run, flaky because prices changed, and about ₹800 a day in API costs across the team.
The rewrite: unit tests with FakeListChatModel, a fake tool-calling model for the agent, and a stubbed price API that returns fixed prices. The suite of 120 tests runs in 5 seconds. One test caught a real bug — when the stock API timed out, the agent used to say "The price is unavailable" and invent a number; the error-path test now asserts no digits appear in that reply. A nightly job runs 40 real-model evaluation cases and posts scores to the team channel.
Follow-up questions to expect
- "How do you fake a model that calls tools?" —
GenericFakeChatModelwith a list ofAIMessages that includetool_calls; withcreate_agentyou may need a small subclass whosebind_toolsreturns itself. - "How do you test prompts?" — Render them with sample inputs and snapshot the result; any change is then visible in review.
- "Can you assert on LLM output at all?" — On real models, use checks that tolerate wording: schema validity, required facts present, LLM-judge scores above a threshold.