Course Content
LangChain Mastery
7 sections · 109 lessons
How do you use LangChain to switch between different LLM providers?
What you need to know
The shared interface is what makes switching cheap. invoke, stream, batch, bind_tools and with_structured_output look the same on ChatOpenAI, ChatAnthropic or ChatGoogleGenerativeAI.
Three ways to switch
1import os2from langchain.chat_models import init_chat_model34# 1. From configuration, at start-up5llm = init_chat_model(os.environ["CHAT_MODEL"], temperature=0)6# CHAT_MODEL="openai:gpt-5.4-mini" or "anthropic:claude-sonnet-4-6"78# 2. Per request, with a configurable model9flex = init_chat_model(configurable_fields=("model", "model_provider"),10 temperature=0)11chain = prompt | flex | parser12chain.invoke(inputs, config={"configurable": {"model": "claude-sonnet-4-6",13 "model_provider": "anthropic"}})1415# 3. Automatic failover16primary = init_chat_model("openai:gpt-5.4-mini", timeout=20)17backup = init_chat_model("anthropic:claude-sonnet-4-6", timeout=20)18robust = primary.with_fallbacks([backup])- Configuration is the simplest and most common. Different environments or customers get different models with no code change.
- Configurable fields let one running service pick the model per call, which is useful for A/B tests or a "premium" tier.
configurable_alternativesdoes the same with a named set of pre-built models. - Fallbacks try the next model only when the first one raises an error.
Model names in these examples are placeholders; read the current names from each provider's model list.
What does not transfer
| Area | Why it can break |
|---|---|
| Prompt wording | A prompt tuned on one model often scores a few points lower on another |
| Tool calling | Models differ in how reliably they pick tools and fill arguments |
| Structured output | Supported methods differ (json_schema, function calling) |
| Context length and price | A long RAG prompt may fit one model and not another |
| Response metadata | finish_reason versus stop_reason, different usage fields |
A real-life example
A support bot over an insurance company's help-centre docs ran on one provider. After a price change, the team wanted to test a cheaper model. Because the model came from init_chat_model(os.environ["CHAT_MODEL"]), the switch in staging was one environment variable. Their evaluation set of 300 real customer questions showed answer accuracy fell from 91% to 86%, mostly on questions that needed two articles combined. They kept the cheaper model for the 70% of traffic classified as simple FAQs and routed the rest to the stronger model, cutting the monthly bill by about 40% without a visible drop in quality.
Follow-up questions to expect
- "How do you make sure fallbacks trigger only on real failures?" — Fallbacks run only when an exception is raised; set timeouts so a slow provider raises instead of hanging, and pass
exceptions_to_handleto limit which errors switch over. - "Can the fallback chain use a different prompt?" — Yes: put
with_fallbackson the whole chain,chain_a.with_fallbacks([chain_b]), wherechain_bhas a prompt tuned for the backup model. - "Why not always use the cheapest model?" — Quality differs by task; measure on your own data and route by difficulty.