LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you ensure reproducibility in LangChain applications?


What you need to know

Why outputs change

  • Sampling — with temperature above 0, the model picks among likely tokens at random.
  • Hidden nondeterminism — even at temperature 0, batching and hardware effects on the provider's side can change outputs slightly.
  • Model updates — an alias like gpt-5 or latest may point to a new snapshot next month.
  • Your own changes — prompt edits, library upgrades, a re-built index with different chunks.

The controls

Source of changeControl
Model versionDated snapshot id in config, never a floating alias
Samplingtemperature=0 (some reasoning models don't accept it); seed where the provider supports it
PromptFiles in git or a prompt registry; log version or hash
LibrariesLockfile with exact versions
RetrievalRecord embedding model, chunk settings, index build id
TestsFake models or recorded responses; a cache in test runs

Record every run

Python
import hashlibprompt_text = open("prompts/answer_v7.txt").read()prompt_hash = hashlib.sha256(prompt_text.encode()).hexdigest()[:12]result = chain.invoke(inputs, config={"metadata": {    "model": settings.chat_model,           # e.g. a dated snapshot id    "temperature": settings.temperature,    "prompt_version": "answer_v7",    "prompt_hash": prompt_hash,    "index_build": settings.index_build_id,    "git_sha": settings.git_sha,}})

The metadata goes to every trace, so any past answer can be traced to the exact model, prompt and index that produced it.

Replay in tests

For unit tests, use fake models (FakeListChatModel) so results are fixed. For integration tests, a persistent cache (set_llm_cache(SQLiteCache(...))) replays recorded model answers, as long as prompts and parameters haven't changed.

Evaluation instead of equality

For quality, run the same dataset through the old and new version and compare scores. If accuracy stays within a point or two, the change is safe even though exact wording differs.

A real-life example

A loan-support assistant at a bank must explain, months later, why it told a customer a certain EMI amount. An auditor asks about a chat from March.

Because each trace carries model snapshot id, prompt hash, index build id and retrieved chunk ids, the team finds the exact policy chunk used and the prompt version. They re-run the same inputs with the same model snapshot and prompt in a notebook and get an answer with the same figures and slightly different wording — enough to confirm the reply followed policy. Before this logging existed, a similar question in the previous year took two weeks and ended with "we cannot tell".

Follow-up questions to expect

  • "Does temperature 0 make it deterministic?" — Close, not guaranteed; that is why you log and evaluate rather than rely on equality.
  • "What if the provider retires your snapshot?" — Treat it like any upgrade: run the eval set on the replacement, compare, then switch.
  • "How do you reproduce an agent run?" — Log tool inputs and outputs too; replay with the tools stubbed to return the recorded results.