LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

What is caching in LangChain, and how is it used?


Exact cache against semantic cacheExact LLM cache• Key is the exact prompt plus params• Never returns a wrong answer• Misses on any change in wording• Great for tests and batch rerunsSemantic cache• Key is embedding similarity• Paraphrases hit the cache• Loose threshold flips the meaning• Only for FAQ-style questions
A semantic cache trades safety for hit rate — 'within 30 days' and 'after 30 days' are close vectors with opposite answers.

What you need to know

Four different caches

CacheWhat it storesHit whenTool
Exact LLM cacheFull model responseSame prompt and same parametersInMemoryCache, SQLiteCache, Redis cache
Semantic cacheFull model responseA similar prompt (embedding distance under a threshold)RedisSemanticCache and similar
Embedding cacheVectors for textsSame text embedded againCacheBackedEmbeddings
Provider prompt cacheProcessed prompt prefix, on the provider's sideSame long prefix (system prompt, documents)Provider feature; no LangChain setup

Exact caching

Python
from langchain_core.globals import set_llm_cachefrom langchain_core.caches import InMemoryCachefrom langchain_community.cache import SQLiteCacheset_llm_cache(InMemoryCache())                          # dev: per processset_llm_cache(SQLiteCache(database_path=".llm_cache.db"))  # persists across runsllm.invoke("What is the refund window?")   # calls the providerllm.invoke("What is the refund window?")   # served from cache

You can also scope a cache to one model: ChatOpenAI(model=..., cache=InMemoryCache()), or turn it off for one model with cache=False.

Semantic caching and its risk

A semantic cache embeds the prompt and returns a stored answer if a past prompt was close enough. "How do I return shoes?" can hit "How can I return my shoes?". The danger is a loose threshold: "Can I return shoes after 30 days?" might hit "…within 30 days?" and get the opposite answer. Use it only for FAQ-like questions, with a strict threshold and tests.

Provider prompt caching

Many providers discount repeated long prefixes. Put the stable parts first — system prompt, tool definitions, fixed documents — and the changing parts (the user question) last. Some providers need an explicit cache marker in the message; check the provider's LangChain integration docs.

Rules

  • Never share a cache across users for personalised answers ("What is my leave balance?").
  • Put tool-calling agents' final answers in a cache only when the tool results are stable.
  • Set expiry for anything based on changing data such as prices.

A real-life example

An e-commerce support bot gets 200,000 questions a day during a sale. About 30% are the same 50 questions ("When will my refund arrive?", "How do I cancel?"), asked with slightly different words.

The team adds an answer cache in front of the RAG chain for FAQ-type questions only (detected by a classifier), with a semantic threshold tuned on 500 labelled pairs, and a 24-hour expiry. Cache hit rate is 22%, those answers return in 40 ms instead of 2.5 s, and model spend falls by about a fifth. Order-specific questions ("Where is order 4431?") skip the cache entirely, because a shared cache would leak one customer's data to another.

Follow-up questions to expect

  • "Why did my cache not hit?" — Any difference in prompt text, model name, temperature or other parameters changes the key — including a timestamp in the system prompt.
  • "Does caching work with streaming?" — A cache hit returns the stored result at once; the stream then arrives as a single chunk.
  • "How do you cache agent tool results?" — Cache inside the tool (for example, price lookups for 60 seconds), which is often more useful than caching the whole agent reply.