Course Content
LangChain Mastery
7 sections · 109 lessons
How do you implement circuit breakers in LangChain applications?
What you need to know
Why retries are not enough
Retries help with short blips. During a real outage, every request still waits for its timeouts and retries — say 3 attempts × 20 seconds — before failing. Threads and connections pile up, users wait a minute for an error, and your retries add load to a provider that is already down. A breaker notices the pattern and stops trying for a while.
The three states
- Closed — calls go through. Count consecutive failures.
- Open — after
thresholdfailures, reject every call at once forcooldownseconds. - Half-open — after the cooldown, let a call through as a probe. Success closes the circuit; failure opens it again.
A breaker as a runnable wrapper
1import threading, time2from langchain_core.runnables import RunnableLambda34class CircuitOpen(Exception):5 pass67class CircuitBreaker:8 def __init__(self, threshold: int = 5, cooldown: float = 30.0):9 self.threshold, self.cooldown = threshold, cooldown10 self.failures, self.opened_at = 0, None11 self.lock = threading.Lock()1213 def wrap(self, runnable):14 def call(inputs, config):15 with self.lock:16 if self.opened_at and time.monotonic() - self.opened_at < self.cooldown:17 raise CircuitOpen("provider circuit is open")18 try:19 result = runnable.invoke(inputs, config) # also the half-open probe20 except Exception:21 with self.lock:22 self.failures += 123 if self.failures >= self.threshold:24 self.opened_at = time.monotonic()25 raise26 with self.lock:27 self.failures, self.opened_at = 0, None28 return result29 return RunnableLambda(call)3031primary = CircuitBreaker(threshold=5, cooldown=30).wrap(prompt | primary_llm)32chain = primary.with_fallbacks([prompt | backup_llm]) | parserWhen the circuit is open, CircuitOpen is raised in microseconds, and with_fallbacks sends the request to the backup model. This simple version lets every request after the cooldown act as a probe; a production breaker allows exactly one.
Production details
- Shared state — a per-process breaker in a 20-pod deployment has 20 separate views. Keep the failure count and open time in Redis, or use a library or service mesh that does this.
- What counts as failure — timeouts, 5xx and connection errors. Not 400s or validation errors; those are your bugs, not the provider's.
- Per dependency — one breaker per provider or tool, so a broken stock API doesn't stop the model.
- Alerting — an open circuit should alert someone.
In agents, ModelFallbackMiddleware handles the fallback part; the breaker can be a custom wrap_model_call middleware using the same logic.
A real-life example
A finance assistant calls a third-party stock-price API as a tool. One morning the API started timing out after 15 seconds on every call. With 2 retries, each question took 45 seconds to fail, worker threads were exhausted, and even questions that didn't need prices stalled.
With a breaker (5 failures, 60-second cooldown) around the tool, after the first five timeouts the tool returned "live prices unavailable" instantly. The agent answered with the last cached close price and a clear note. Other questions stayed fast. The breaker's "open" event paged the on-call engineer within a minute, and a probe every 60 seconds closed the circuit automatically when the API recovered 25 minutes later.
Follow-up questions to expect
- "Breaker vs retry — which first?" — Retries inside, breaker outside: retries absorb blips; the breaker stops retry storms during outages.
- "How do you choose threshold and cooldown?" — From the provider's normal error rate and typical outage length; start with 5 failures and 30–60 seconds, then tune.
- "Rate limiter or breaker?" — A rate limiter (such as
InMemoryRateLimiteron the chat model) prevents you from causing 429s; a breaker reacts when the dependency is failing anyway.