Course Content
LangChain Mastery
7 sections · 109 lessons
How do you ensure reproducibility in LangChain applications?
What you need to know
Why outputs change
- Sampling — with temperature above 0, the model picks among likely tokens at random.
- Hidden nondeterminism — even at temperature 0, batching and hardware effects on the provider's side can change outputs slightly.
- Model updates — an alias like
gpt-5orlatestmay point to a new snapshot next month. - Your own changes — prompt edits, library upgrades, a re-built index with different chunks.
The controls
| Source of change | Control |
|---|---|
| Model version | Dated snapshot id in config, never a floating alias |
| Sampling | temperature=0 (some reasoning models don't accept it); seed where the provider supports it |
| Prompt | Files in git or a prompt registry; log version or hash |
| Libraries | Lockfile with exact versions |
| Retrieval | Record embedding model, chunk settings, index build id |
| Tests | Fake models or recorded responses; a cache in test runs |
Record every run
1import hashlib23prompt_text = open("prompts/answer_v7.txt").read()4prompt_hash = hashlib.sha256(prompt_text.encode()).hexdigest()[:12]56result = chain.invoke(inputs, config={"metadata": {7 "model": settings.chat_model, # e.g. a dated snapshot id8 "temperature": settings.temperature,9 "prompt_version": "answer_v7",10 "prompt_hash": prompt_hash,11 "index_build": settings.index_build_id,12 "git_sha": settings.git_sha,13}})The metadata goes to every trace, so any past answer can be traced to the exact model, prompt and index that produced it.
Replay in tests
For unit tests, use fake models (FakeListChatModel) so results are fixed. For integration tests, a persistent cache (set_llm_cache(SQLiteCache(...))) replays recorded model answers, as long as prompts and parameters haven't changed.
Evaluation instead of equality
For quality, run the same dataset through the old and new version and compare scores. If accuracy stays within a point or two, the change is safe even though exact wording differs.
A real-life example
A loan-support assistant at a bank must explain, months later, why it told a customer a certain EMI amount. An auditor asks about a chat from March.
Because each trace carries model snapshot id, prompt hash, index build id and retrieved chunk ids, the team finds the exact policy chunk used and the prompt version. They re-run the same inputs with the same model snapshot and prompt in a notebook and get an answer with the same figures and slightly different wording — enough to confirm the reply followed policy. Before this logging existed, a similar question in the previous year took two weeks and ended with "we cannot tell".
Follow-up questions to expect
- "Does temperature 0 make it deterministic?" — Close, not guaranteed; that is why you log and evaluate rather than rely on equality.
- "What if the provider retires your snapshot?" — Treat it like any upgrade: run the eval set on the replacement, compare, then switch.
- "How do you reproduce an agent run?" — Log tool inputs and outputs too; replay with the tools stubbed to return the recorded results.