Course Content
Building AI Features in React
5 sections · 21 lessons
Cost and Rate Limits From the Client
ShopLens's first full week in production cost $410 a day. The estimate had been $60. Three causes explained almost all of the difference. The summary was not cached yet, so every product view generated a new one. One browser tab made 1,100 compare calls in 20 minutes, because an effect depended on an aspects array that was a new object on every render. And search-engine crawlers that execute JavaScript were loading product pages and triggering summaries nobody would read.
All three are frontend or BFF decisions. The model's price did not change. This lesson goes through the tools that brought the cost down: arithmetic first, then caching, deduplication and limits.
Know the numbers before you optimise
Start with a table of traffic and cost per call, using the prices from section 1. It tells you where to spend your effort.
| Feature | Calls per day, before | Cost per call | Cost per day, before | After this lesson |
|---|---|---|---|---|
| Summary | 16,000 (every view that reached the panel) | $0.02 | $320 | about $30 (one per product per day, plus regenerates) |
| Compare | 2,400 (including one runaway tab) | $0.03 | $72 | about $6 (cached pairs, no loop) |
| Badge | 4,500 | $0.004 | $18 | about $2 (labels stored at write time) |
| Total | $410 | about $38 |
The summary dominates, so caching it is the first job. This is almost always the pattern: one feature is most of the bill, and one design change fixes most of it.
Cache on the server, keyed by what changes the answer
A cache key must change when, and only when, the answer should change. For the summary, three things change the answer: the product, the reviews that go into the prompt and the prompt itself. So the key contains all three.
1// server/routes/summary.ts (caching added)2import { createHash } from "node:crypto";3import type { Review } from "../../shared/schemas";4import { cache } from "../cache"; // get/set with a TTL; Redis in production5import { SUMMARY_PROMPT_VERSION } from "../prompts";67const DAY_MS = 24 * 60 * 60 * 1000;89function summaryKey(productId: string, reviews: Review[]): string {10 const inputs = createHash("sha256").update(reviews.map((r) => r.id).join(",")).digest("hex");11 return `summary:${SUMMARY_PROMPT_VERSION}:${productId}:${inputs.slice(0, 16)}`;12}1314// Inside summaryRoute, after loading the product and its reviews:15const key = summaryKey(productId, reviews);16const hit = input.data.fresh ? undefined : await cache.get(key);17if (hit) {18 res.type("text/plain").send(hit); // about 30 ms instead of about 3 s19 return;20}21// ...stream from the model as before; pipeSummary now also returns the text it sent...22if (result.outcome === "sent") await cache.set(key, result.text, DAY_MS);The key uses a hash of the review ids that were actually sent, so a new review that enters the top 40 produces a new key, and a new review that does not enter it changes nothing. The prompt version is in the key, so deploying summary-v4 naturally starts fresh summaries. Only complete summaries are stored: a stopped or cut stream is never cached, so one shopper's Stop never becomes everyone's summary.
The compare cache uses the same idea, with one twist. "Kestrel X2 against Wren Air" and "Wren Air against Kestrel X2" are the same question, so the key uses the two ids sorted. When the request's order differs from the cached order, the route swaps the columns:
1// server/routes/compare.ts (addition)2const FLIP = { a: "b", b: "a", tie: "tie", unclear: "unclear" } as const;34function swapColumns(reply: CompareReply): CompareReply {5 return { rows: reply.rows.map((r) => ({ ...r, a: r.b, b: r.a, better: FLIP[r.better] })) };6}For the most popular products, go one step further and generate the summary in a background job when reviews change, like the badge labels in section 3. Then the page path is nearly always a cache hit, and 50 shoppers arriving at once from a newsletter never trigger 50 simultaneous generations.
In the browser: do not ask twice
The browser's job is to avoid sending requests that are not needed. Four habits cover most cases.
- Stable dependencies. The runaway tab was caused by
useCompare(productId, otherId, aspects)withaspectscreated inline. Depend onaspects.join("|")instead, as the badge hook does with its ids. - Explicit triggers for edits. Aspect chips update on "Update comparison", not on each toggle. For fields that must react while typing, debounce by 300 to 500 ms and abort the previous request.
- Only what is visible. Start the summary when the panel is 300 pixels from the viewport (section 4), so pages that are never scrolled cost nothing. Most crawlers never scroll.
- Remember answers for the session. A shopper who compares Kestrel X2 with Wren Air, then another model, then Wren Air again should get the third answer instantly.
The last habit is a small promise cache. Caching the promise, not the result, also deduplicates: two components asking the same question at the same moment share one request.
1// src/api/cache.ts2const results = new Map<string, Promise<unknown>>();34export function cached<T>(key: string, load: () => Promise<T>): Promise<T> {5 const hit = results.get(key) as Promise<T> | undefined;6 if (hit) return hit;7 const promise = load().catch((err: unknown) => {8 results.delete(key); // never cache a failure9 throw err;10 });11 results.set(key, promise);12 return promise;13}useCompare now calls cached(JSON.stringify(body), () => withRetry(() => postJson("/api/compare", body, CompareReply, signal), signal)), where signal is AbortSignal.timeout(30_000). Notice what changed: the component's own abort signal is no longer passed in, because a shared request must not be cancelled by the first component that stops caring. Instead the effect sets a local current = false flag in its cleanup and ignores late results. This is a deliberate trade-off. Compare results are cheap to keep and likely to be reused, so letting them finish is worth it. The summary stream keeps its end-to-end abort, because a stopped stream is not reused.
Rate limits: yours and the provider's
Caching controls cost from normal traffic. Rate limits control cost from abnormal traffic: bugs, scripts and abuse. ShopLens has three layers:
1// server/app.ts (addition, before the routes)2import { rateLimit } from "express-rate-limit";34app.use("/api", rateLimit({ windowMs: 60_000, limit: 60, standardHeaders: "draft-7", legacyHeaders: false }));5app.use("/api/compare", rateLimit({ windowMs: 60 * 60_000, limit: 30, standardHeaders: "draft-7", legacyHeaders: false }));The first line allows 60 AI requests a minute from one client, far more than a real shopper makes. The second allows 30 comparisons an hour. Regenerate has its own limit of two per product per session, checked in the summary route. These limits send a 429 with a Retry-After header, which the browser client from section 3 already reads, and which section 4 turned into a visible countdown.
The provider has rate limits too, on requests and tokens per minute. When your BFF receives a 429 from the provider, pass it on as a 429 with a Retry-After value, rather than a generic 502, so the browser waits instead of retrying at once. A provider 429 during normal traffic is also a signal to review your caching: most of the time it means the same work is being done many times.
Check your understanding
0 of 3 answered
1.Which cache key is right for the summary?
2.Why does useCompare stop passing its own abort signal into the shared, cached request?
3.One browser tab made 1,100 compare calls in 20 minutes. What was the cause?