Building with LLMs

Handling Latency, Retries, and Caching


At 14:02 on a Thursday, a model provider returns 503 to about one request in twenty. It is a minor blip; their status page later describes it as eleven minutes of degraded service.

Your service is down for fifty minutes.

The reason is a loop somebody wrote in an afternoon:

Python
for attempt in range(5):    try:        return client.chat.completions.create(**kwargs)    except Exception:        continue          # try again immediately

When 5% of calls started failing, every failing request instantly became five requests. Your outbound traffic to the provider jumped, which pushed you into rate limiting, which produced 429s, which the same loop also retried immediately, which produced more 429s. By 14:06 essentially every request was failing and retrying in a tight loop, consuming your worker threads. The provider recovered at 14:13. Your service did not recover until someone deployed a fix at 14:52, because your own clients were still hammering it.

Nothing exotic happened. A retry loop with no delay, no jitter and no discrimination between error types turned a 5% upstream error rate into a total outage. This lesson is about the four things that stop that: retrying correctly, respecting rate limits before you hit them, caching what does not need recomputing, and getting bytes to the user faster than the model finishes thinking.

Backoff with jitter, in seconds0.51.02.04.08.0stop012345first retryopencircuitEach wait is multiplied by a random factor, so a thousand clients do not retry in lockstep.
Doubling spreads the load over time; jitter spreads it across callers — without both, the retry storm is the outage.

Retrying without amplifying the failure

Retry the right errors only

The first bug in that loop is except Exception. Most errors will not succeed on a second attempt, and retrying them wastes time and money while hiding the real cause.

ErrorRetry?Why
429 rate limitYes — after the delay the server asks forTransient by definition
500, 502, 503, 504Yes, with backoffServer-side, usually brief
Connection error, timeoutYes, with backoffNetwork conditions change
400 bad requestNoYour request is malformed; it will be malformed again
401 / 403NoCredentials will not fix themselves
404 model not foundNoTypo in the model name
Context length exceededNo — shrink the input insteadSame input, same failure
Content filter refusalNoDeterministic policy decision

Exponential backoff, and why jitter is not optional

Backoff spaces retries out so a struggling service gets room to recover. Jitter randomises the spacing so a thousand clients do not all retry at the same instant.

Python
import random, time, loggingfrom openai import RateLimitError, APITimeoutError, APIConnectionError, InternalServerErrorRETRYABLE = (RateLimitError, APITimeoutError, APIConnectionError, InternalServerError)def with_backoff(fn, attempts=5, base=1.0, cap=32.0):    for i in range(attempts):        try:            return fn()        except RETRYABLE as exc:            if i == attempts - 1:                logging.error("giving up after %d attempts: %s", attempts, exc)                raise            response = getattr(exc, "response", None)      # absent on timeouts            retry_after = response.headers.get("retry-after") if response is not None else None            if retry_after:                delay = float(retry_after)            else:                delay = min(cap, base * (2 ** i)) * (0.5 + random.random())            logging.warning("attempt %d failed (%s); sleeping %.2fs",                            i + 1, type(exc).__name__, delay)            time.sleep(delay)

Work out what those settings actually cost you. Nominal delays are 1, 2, 4, 8 seconds before jitter; with the 0.5×–1.5× multiplier each falls somewhere in 0.5–1.5, 1–3, 2–6 and 4–12 seconds. Worst case total sleep is about 22.5 seconds, on top of four failed request attempts. If your HTTP handler has a 30-second timeout, this configuration can consume the entire budget and return nothing.

Compute that number for your own settings, always. A retry policy that cannot complete inside your request timeout is a policy that guarantees timeouts under load.

Include the retries you did not write. The OpenAI and Anthropic Python clients already retry connection errors, 429s and 5xx responses twice by default, with their own backoff. Wrap them in a five-attempt loop and one failing request can become fifteen calls. When you own the retry policy, construct the client with max_retries=0 (for example OpenAI(max_retries=0)); when you do not need a custom policy, the built-in one is a sensible default.

Now the jitter argument, with numbers. Suppose 1,000 clients hit a 503 at the same moment.

Without jitterWith jitter (0.5×–1.5×)
At t = 1 s1,000 simultaneous retries~1,000 spread over 0.5–1.5 s ≈ 1,000/s peak
At t = 2 s1,000 simultaneous retriesSpread over 1–3 s ≈ 500/s peak
At t = 4 s1,000 simultaneous retriesSpread over 2–6 s ≈ 250/s peak
Effect on recovering serviceRepeated full-load spikesLoad flattens and decays

Without jitter, the retries synchronise into waves, and each wave re-breaks a service that was about to recover. The waves get no smaller with time; they just get further apart. Jitter is what converts a spike into a slope.

Retries without jitter do not spread load — they synchronise it. Your own retry logic is then a denial-of-service attack that fires precisely when the target is weakest.

The circuit breaker: stop trying when it is clearly down

Backoff still means every request pays the full retry cost during a sustained outage. A circuit breaker notices the pattern and fails fast instead.

Python
import time, threadingclass CircuitBreaker:    def __init__(self, threshold=5, cooldown=60):        self.threshold, self.cooldown = threshold, cooldown        self.failures, self.opened_at = 0, None        self.lock = threading.Lock()    def call(self, fn):        with self.lock:            if self.opened_at and time.time() - self.opened_at < self.cooldown:                raise RuntimeError("circuit open: upstream unavailable")            if self.opened_at:                      # cooldown elapsed — try one probe                self.opened_at = None        try:            result = fn()        except Exception:            with self.lock:                self.failures += 1                if self.failures >= self.threshold:                    self.opened_at = time.time()            raise        with self.lock:            self.failures = 0        return result

The payoff is in latency, not just load. With a 22-second retry budget and a fully down provider, every user waits 22 seconds for an error. With the breaker open, they get an error in under a millisecond and your service can show a useful message or fall back. Failing fast is a feature.

Rate limits: stay under them rather than bouncing off

Reacting to 429s is strictly worse than never causing them, because every 429 is a request you paid latency for and got nothing from. Two controls do almost all the work.

Bound concurrency

Python
import asyncioMAX_CONCURRENT = 8sem = asyncio.Semaphore(MAX_CONCURRENT)async def call_model(payload):    async with sem:        return await async_client.chat.completions.create(**payload)

Without this, asyncio.gather over 500 documents opens 500 simultaneous connections and you are rate limited within a second. The right number is derived, not guessed. If a call takes 2 seconds, one worker completes 60 / 2 = 30 requests per minute, so n workers produce 30n. Against a limit of 500 requests per minute, 30n ≤ 500 gives n ≤ 16. Pick 12, to leave headroom for other traffic on the same key.

Spend the token budget, not just the request budget

Most providers limit tokens per minute as well as requests per minute, and the token limit is usually the one you hit first. A token bucket enforces it locally:

Python
import time, threadingclass TokenBucket:    def __init__(self, tokens_per_minute: int):        self.rate = tokens_per_minute / 60.0        self.capacity = float(tokens_per_minute)        self.tokens = self.capacity        self.updated = time.monotonic()        self.lock = threading.Lock()    def consume(self, amount: int, timeout: float = 60.0):        deadline = time.monotonic() + timeout        while True:            with self.lock:                now = time.monotonic()                self.tokens = min(self.capacity, self.tokens + (now - self.updated) * self.rate)                self.updated = now                if self.tokens >= amount:                    self.tokens -= amount                    return                shortfall = (amount - self.tokens) / self.rate            if time.monotonic() + shortfall > deadline:                raise TimeoutError("token budget not available in time")            time.sleep(min(shortfall, 1.0))bucket = TokenBucket(150_000)def guarded_call(messages, expected_output=800):    bucket.consume(count_tokens(messages) + expected_output)    return client.chat.completions.create(model=MODEL, messages=messages)

Reserve the expected output tokens as well as the input, because the limit counts both. Under-reserving means you drift over the line during bursts, and the 429s arrive exactly when traffic is highest.

Caching: the cheapest optimisation there is

An identical request answered from cache costs nothing and returns in a millisecond instead of two seconds. LLM caching differs from ordinary web caching in three ways worth naming:

  • The key must cover everything that affects the output — model, full prompt, sampling settings, tools, system message. Leave any one of them out and two requests that differ only there share a cached answer, which is worse than a miss.
  • Model output varies from call to call, so caching quietly converts a varied assistant into a repetitive one. You often cannot turn the variation off: Claude Opus 5 and Sonnet 5 reject temperature, and OpenAI's GPT-6 models reject it when reasoning is on. Cache aggressively where one right answer exists (classification, extraction, FAQ answers); think carefully about open-ended replies.
  • The value of a hit is enormous relative to the cost of a miss — you are saving seconds and cents, not milliseconds and bytes. That justifies generous TTLs.

In-process cache

Python
import hashlib, json, timeclass ResponseCache:    def __init__(self, ttl=3600, max_entries=5000):        self.ttl, self.max_entries, self.store = ttl, max_entries, {}        self.hits = self.misses = 0    def _key(self, model, messages, **params) -> str:        blob = json.dumps({"m": model, "msgs": messages, "p": params}, sort_keys=True)        return hashlib.sha256(blob.encode()).hexdigest()    def get_or_call(self, fn, model, messages, **params):        k = self._key(model, messages, **params)        entry = self.store.get(k)        if entry and time.time() - entry[1] < self.ttl:            self.hits += 1            return entry[0]        self.misses += 1        value = fn()        if len(self.store) >= self.max_entries:            oldest = min(self.store, key=lambda kk: self.store[kk][1])            del self.store[oldest]        self.store[k] = (value, time.time())        return value    @property    def hit_rate(self):        total = self.hits + self.misses        return self.hits / total if total else 0.0

sort_keys=True matters more than it looks: without it, two logically identical requests whose dictionaries were built in a different order hash differently and never hit. And max_entries matters because an unbounded cache of model responses is a memory leak with a slow fuse — 50,000 cached answers at 4 KB each is 200 MB.

Shared cache across processes

Python
import redis, json, hashlibr = redis.from_url("redis://localhost:6379/1")def cached_call(fn, model, messages, ttl=3600, **params):    blob = json.dumps({"m": model, "msgs": messages, "p": params}, sort_keys=True)    key = "llm:" + hashlib.sha256(blob.encode()).hexdigest()    hit = r.get(key)    if hit:        return json.loads(hit)    value = fn()    r.setex(key, ttl, json.dumps(value))    return value

With four web workers, an in-process cache gives each one its own copy and your effective hit rate is roughly a quarter of what it could be. Redis shares one cache across all of them, and survives deploys.

What a hit rate is worth

Take 40,000 requests a day averaging 2,000 input and 500 output tokens on a model at 3 and 15 dollars per million. Cost per call is 2,000/1e6 × 3 + 500/1e6 × 15 = 0.006 + 0.0075 = 0.0135 dollars. Uncached daily spend is 40,000 × 0.0135 = 540 dollars.

Hit ratePaid calls/dayDaily costMonthly saving
0%40,000540.00—
15%34,000459.002,430
30%28,000378.004,860
50%20,000270.008,100

All figures in US dollars, monthly saving at 30 days. A 30% hit rate — entirely achievable for FAQ-style traffic — is nearly 4,900 dollars a month from about forty lines of code. Which is why the first thing to measure is your hit rate, and the second is which requests are missing that ought to hit.

Prompt caching: the provider-side variant

Exact-match caching only helps when the whole request repeats. Prompt caching helps when only the prefix repeats — a long system prompt, a fixed set of tool definitions, a document you are asking many questions about. The provider stores the processed prefix and bills subsequent hits at a fraction of the input rate.

It has one rule that governs everything: it is a prefix match, so any byte change invalidates everything after it. Put stable content first and volatile content last. A timestamp, a request ID or an unsorted JSON blob near the top of your system prompt gives you a 0% cache rate and no error message to explain why. If you enable it, check the reported cached-token count on real traffic — a silent invalidator is the normal outcome of a first attempt.

Latency: perceived and actual

Stream, and the wait stops mattering

A 600-token answer takes about four seconds to generate. Non-streaming, the user sees a spinner for four seconds. Streaming, the first words appear in about 400 ms.

Python
stream = client.chat.completions.create(    model=MODEL, messages=messages, stream=True)for chunk in stream:    delta = chunk.choices[0].delta.content    if delta:        yield delta
Non-streamingStreaming
Time to first visible text4.0 s0.4 s
Time to complete answer4.0 s4.1 s
User's experience"Is it broken?""It's working"
AbandonmentHighLow

Total time is marginally worse when streaming. The perceived latency improves tenfold. That trade is almost always right for anything a human is watching — and almost always pointless for a background job, where nobody is watching and streaming only adds complexity.

Two things streaming complicates. You cannot inspect a complete response before showing it, so content filtering has to happen on the fly or not at all. And errors can arrive mid-stream, after you have already displayed 200 tokens, so your UI needs a way to show a failure that interrupts partial output.

Overlap independent work

Python
import asyncioasync def enrich(text):    summary, entities, sentiment = await asyncio.gather(        summarise(text), extract_entities(text), classify_sentiment(text))    return {"summary": summary, "entities": entities, "sentiment": sentiment}

At 2.1, 1.8 and 0.6 seconds, sequential is 4.5 s and concurrent is 2.1 s — a 53% cut for the same three calls and the same cost. The constraint is the concurrency bound from earlier: fanning out is free in money and cheap in time, but it spends your rate limit in a burst.

Pick the smallest model that passes

The largest latency lever is not code. A small-tier model typically responds in a third to a half the time of a large one on the same prompt. Route by task: classification and routing to the small model, hard reasoning to the large one. You get lower latency and lower cost from the same change.

Knowing what you are spending, as you spend it

Every call reports its token usage. Recording it is the difference between managing cost and discovering it.

Python
from dataclasses import dataclass, fieldfrom collections import defaultdictPRICES = {                     # US dollars per 1M tokens    "gpt-6-luna": (0.10, 0.50),        # published rates, September 2026    "claude-sonnet-5": (2.00, 10.00),    "claude-opus-5": (5.00, 25.00),}@dataclassclass Budget:    monthly_limit: float    spend: float = 0.0    by_feature: dict = field(default_factory=lambda: defaultdict(float))    def record(self, model, usage, feature="unknown") -> float:        pin, pout = PRICES[model]        cost = usage.prompt_tokens / 1e6 * pin + usage.completion_tokens / 1e6 * pout        self.spend += cost        self.by_feature[feature] += cost        if self.spend > self.monthly_limit:            raise RuntimeError(f"monthly budget exceeded: {self.spend:.2f}")        if self.spend > self.monthly_limit * 0.8:            logging.warning("80%% of monthly budget used: %.2f", self.spend)        return cost

Tag by feature. A single total tells you the bill rose; a per-feature breakdown tells you that document summarisation is 68% of it and everything else is noise, which is the sentence that leads to a fix.

Log input tokens, output tokens, latency and cache outcome on every single call from the first commit. Every optimisation in this lesson is guesswork without those four numbers, and retrofitting them costs more than adding them.

Where people get this wrong

Retrying everything. The opening story. Discriminate by error type or your retries become the outage.

Backoff without jitter. Synchronises clients into waves and prevents the recovery it was meant to allow.

A retry budget longer than the request timeout. Guarantees timeouts under exactly the conditions retries exist for. Do the addition.

Caching open-ended answers without deciding to. Turns deliberate variety into accidental repetition.

Cache keys that miss a parameter. Leave the model or the system prompt out of the key and you will serve an answer from the wrong configuration. That bug is very hard to find and very easy to prevent.

Unbounded in-memory caches. A memory leak that surfaces as an out-of-memory kill weeks later.

Streaming everything. A batch job that streams gains nothing and adds failure modes.

Optimising before measuring. Teams routinely spend a fortnight on caching for a workload whose real problem was a 4,000-token system prompt sent on every call.

The order to do this in

All of it matters eventually. The sequence matters more, because each step tells you whether the next is worth doing.

  1. Instrument. Tokens in, tokens out, latency, cache hit or miss, model, feature tag, on every call. One day of work; every later decision depends on it.
  2. Fix retries. Discriminate by error type, add exponential backoff with jitter, cap the total retry time below your request timeout, add a circuit breaker. This is the change that prevents an outage rather than saving money, so it comes before the money work.
  3. Bound concurrency and token throughput. Cheap, and it removes the 429s that make everything else look worse than it is.
  4. Stream anything a human waits for. The largest perceived-latency win available, and usually a small code change.
  5. Cache. Exact-match first because it is simple, then prompt caching if you have long stable prefixes. Now you have hit-rate numbers to prove it worked.
  6. Right-size the models. Route by task. Do this last, with your evaluation set in hand, so you can see exactly what quality you traded for the latency and cost.

The team in the opening story eventually did all six. The one that saved them the next time a provider had a bad Thursday was step two, and it took an afternoon.