Course Content
LangChain Mastery
7 sections · 109 lessons
What is caching in LangChain, and how is it used?
What you need to know
Four different caches
| Cache | What it stores | Hit when | Tool |
|---|---|---|---|
| Exact LLM cache | Full model response | Same prompt and same parameters | InMemoryCache, SQLiteCache, Redis cache |
| Semantic cache | Full model response | A similar prompt (embedding distance under a threshold) | RedisSemanticCache and similar |
| Embedding cache | Vectors for texts | Same text embedded again | CacheBackedEmbeddings |
| Provider prompt cache | Processed prompt prefix, on the provider's side | Same long prefix (system prompt, documents) | Provider feature; no LangChain setup |
Exact caching
1from langchain_core.globals import set_llm_cache2from langchain_core.caches import InMemoryCache3from langchain_community.cache import SQLiteCache45set_llm_cache(InMemoryCache()) # dev: per process6set_llm_cache(SQLiteCache(database_path=".llm_cache.db")) # persists across runs78llm.invoke("What is the refund window?") # calls the provider9llm.invoke("What is the refund window?") # served from cacheYou can also scope a cache to one model: ChatOpenAI(model=..., cache=InMemoryCache()), or turn it off for one model with cache=False.
Semantic caching and its risk
A semantic cache embeds the prompt and returns a stored answer if a past prompt was close enough. "How do I return shoes?" can hit "How can I return my shoes?". The danger is a loose threshold: "Can I return shoes after 30 days?" might hit "…within 30 days?" and get the opposite answer. Use it only for FAQ-like questions, with a strict threshold and tests.
Provider prompt caching
Many providers discount repeated long prefixes. Put the stable parts first — system prompt, tool definitions, fixed documents — and the changing parts (the user question) last. Some providers need an explicit cache marker in the message; check the provider's LangChain integration docs.
Rules
- Never share a cache across users for personalised answers ("What is my leave balance?").
- Put tool-calling agents' final answers in a cache only when the tool results are stable.
- Set expiry for anything based on changing data such as prices.
A real-life example
An e-commerce support bot gets 200,000 questions a day during a sale. About 30% are the same 50 questions ("When will my refund arrive?", "How do I cancel?"), asked with slightly different words.
The team adds an answer cache in front of the RAG chain for FAQ-type questions only (detected by a classifier), with a semantic threshold tuned on 500 labelled pairs, and a 24-hour expiry. Cache hit rate is 22%, those answers return in 40 ms instead of 2.5 s, and model spend falls by about a fifth. Order-specific questions ("Where is order 4431?") skip the cache entirely, because a shared cache would leak one customer's data to another.
Follow-up questions to expect
- "Why did my cache not hit?" — Any difference in prompt text, model name, temperature or other parameters changes the key — including a timestamp in the system prompt.
- "Does caching work with streaming?" — A cache hit returns the stored result at once; the stream then arrives as a single chunk.
- "How do you cache agent tool results?" — Cache inside the tool (for example, price lookups for 60 seconds), which is often more useful than caching the whole agent reply.