LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement unit tests for LangChain chains?


Where each kind of test runsUnit: fakemodel, every commitUnit: spy on therendered promptIntegration:real calls, nightlyEval: LangSmithdataset, scoredtopbottom120 unit tests with fakes run in about 5 seconds.
Fakes make the logic tests fast and exact; quality is judged separately, with scores instead of string equality.

What you need to know

The testing layers

LayerModelChecksWhen
UnitFakePrompt rendering, parsing, routing, filters, error handlingEvery commit, seconds
IntegrationReal, few callsProvider wiring, structured output with the real modelNightly or before release
EvaluationReal, datasetQuality: correctness, groundedness, tool choiceOn prompt/model/retrieval changes

A unit test with fakes

Python
from langchain_core.documents import Documentfrom langchain_core.language_models.fake_chat_models import FakeListChatModelfrom langchain_core.runnables import RunnableLambdadef test_rag_puts_retrieved_text_in_prompt():    seen = {}    def spy(prompt_value):                 # records the prompt sent to the model        seen["prompt"] = prompt_value.to_string()        return prompt_value    llm = FakeListChatModel(responses=["You get 30 days [refunds.md]"])    retriever = RunnableLambda(lambda q: [        Document(page_content="Refunds within 30 days.", metadata={"source": "refunds.md"})])    chain = build_rag_chain(retriever, RunnableLambda(spy) | llm)    out = chain.invoke("What is the refund window?")    assert "Refunds within 30 days." in seen["prompt"]    assert "[refunds.md]" in out["answer"]    assert out["context"][0].metadata["source"] == "refunds.md"
  • The fake model returns its responses in order, with no network.
  • The spy step checks what the model would have received — the most important thing in a RAG test.
  • Assertions are about structure (source present, context passed), not wording.

What else to unit test

  • Parsers and validation — feed malformed text or objects; assert a clean error or repair.
  • Tools — call the tool function directly with valid and invalid arguments (tool.invoke({...})); mock the external API.
  • Agents — use a fake chat model that returns an AIMessage with tool_calls, then assert the right tool ran with the right arguments and that its result reached the next model turn.
  • Security — the retriever for tenant A never returns tenant B's documents.
  • Error paths — make a fake raise TimeoutError and assert the fallback message.

Quality in CI

With pip install "langsmith[pytest]", tests marked @pytest.mark.langsmith log inputs, outputs and feedback to LangSmith, so an evaluation suite can run as ordinary pytest with results tracked over time.

A real-life example

A finance assistant has a get_stock_price tool and an answer chain. Its tests used to call the real model and the real stock API: 4 minutes per run, flaky because prices changed, and about ₹800 a day in API costs across the team.

The rewrite: unit tests with FakeListChatModel, a fake tool-calling model for the agent, and a stubbed price API that returns fixed prices. The suite of 120 tests runs in 5 seconds. One test caught a real bug — when the stock API timed out, the agent used to say "The price is unavailable" and invent a number; the error-path test now asserts no digits appear in that reply. A nightly job runs 40 real-model evaluation cases and posts scores to the team channel.

Follow-up questions to expect

  • "How do you fake a model that calls tools?" — GenericFakeChatModel with a list of AIMessages that include tool_calls; with create_agent you may need a small subclass whose bind_tools returns itself.
  • "How do you test prompts?" — Render them with sample inputs and snapshot the result; any change is then visible in review.
  • "Can you assert on LLM output at all?" — On real models, use checks that tolerate wording: schema validity, required facts present, LLM-judge scores above a threshold.