LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you implement circuit breakers in LangChain applications?


The three states of a breaker around a stock APIClosed:calls pass,failures countedOpen after 5failures:reject instantlyHalf-openafter 60 s:one probe callProbesucceeds:closed againWhile open, with_fallbacks answers from cache or a second provider.
Retries absorb blips; the breaker stops a 45-second failure from being repeated on every request during an outage.

What you need to know

Why retries are not enough

Retries help with short blips. During a real outage, every request still waits for its timeouts and retries — say 3 attempts × 20 seconds — before failing. Threads and connections pile up, users wait a minute for an error, and your retries add load to a provider that is already down. A breaker notices the pattern and stops trying for a while.

The three states

  1. Closed — calls go through. Count consecutive failures.
  2. Open — after threshold failures, reject every call at once for cooldown seconds.
  3. Half-open — after the cooldown, let a call through as a probe. Success closes the circuit; failure opens it again.

A breaker as a runnable wrapper

Python
import threading, timefrom langchain_core.runnables import RunnableLambdaclass CircuitOpen(Exception):    passclass CircuitBreaker:    def __init__(self, threshold: int = 5, cooldown: float = 30.0):        self.threshold, self.cooldown = threshold, cooldown        self.failures, self.opened_at = 0, None        self.lock = threading.Lock()    def wrap(self, runnable):        def call(inputs, config):            with self.lock:                if self.opened_at and time.monotonic() - self.opened_at < self.cooldown:                    raise CircuitOpen("provider circuit is open")            try:                result = runnable.invoke(inputs, config)      # also the half-open probe            except Exception:                with self.lock:                    self.failures += 1                    if self.failures >= self.threshold:                        self.opened_at = time.monotonic()                raise            with self.lock:                self.failures, self.opened_at = 0, None            return result        return RunnableLambda(call)primary = CircuitBreaker(threshold=5, cooldown=30).wrap(prompt | primary_llm)chain = primary.with_fallbacks([prompt | backup_llm]) | parser

When the circuit is open, CircuitOpen is raised in microseconds, and with_fallbacks sends the request to the backup model. This simple version lets every request after the cooldown act as a probe; a production breaker allows exactly one.

Production details

  • Shared state — a per-process breaker in a 20-pod deployment has 20 separate views. Keep the failure count and open time in Redis, or use a library or service mesh that does this.
  • What counts as failure — timeouts, 5xx and connection errors. Not 400s or validation errors; those are your bugs, not the provider's.
  • Per dependency — one breaker per provider or tool, so a broken stock API doesn't stop the model.
  • Alerting — an open circuit should alert someone.

In agents, ModelFallbackMiddleware handles the fallback part; the breaker can be a custom wrap_model_call middleware using the same logic.

A real-life example

A finance assistant calls a third-party stock-price API as a tool. One morning the API started timing out after 15 seconds on every call. With 2 retries, each question took 45 seconds to fail, worker threads were exhausted, and even questions that didn't need prices stalled.

With a breaker (5 failures, 60-second cooldown) around the tool, after the first five timeouts the tool returned "live prices unavailable" instantly. The agent answered with the last cached close price and a clear note. Other questions stayed fast. The breaker's "open" event paged the on-call engineer within a minute, and a probe every 60 seconds closed the circuit automatically when the API recovered 25 minutes later.

Follow-up questions to expect

  • "Breaker vs retry — which first?" — Retries inside, breaker outside: retries absorb blips; the breaker stops retry storms during outages.
  • "How do you choose threshold and cooldown?" — From the provider's normal error rate and typical outage length; start with 5 failures and 30–60 seconds, then tune.
  • "Rate limiter or breaker?" — A rate limiter (such as InMemoryRateLimiter on the chat model) prevents you from causing 429s; a breaker reacts when the dependency is failing anyway.