Building AI Features in React

The Model From the Browser's Point of View


Several model providers publish JavaScript SDKs that can run in the browser. One of them makes you pass an option called dangerouslyAllowBrowser: true before it will do so. The name is honest. Calling a model directly from the browser is fine for a local prototype on your own machine and wrong for anything a stranger can open.

ShopLens therefore has a small server of its own between the browser and the model. This pattern is called a backend-for-frontend, or BFF: a server that exists to serve one frontend, shaped around that frontend's needs. This lesson explains what goes wrong without one, what the BFF must own, and how to design its endpoints so they do not become a free model for the internet.

What the backend-for-frontend ownsShopLens BFFThe API key, never bundledThe prompt and its versionReview data, loaded by idToken, word and rate limitsValidationbefore output leavesCaching and model choice
Hiding the key is only step one — a generic prompt proxy still gives anyone a free model on your bill.

Why the key cannot live in the browser

Any value your frontend build can read ends up in the JavaScript bundle. In Vite, every environment variable that starts with VITE_ is inlined into the code at build time. In Next.js, the same is true for NEXT_PUBLIC_. Anyone can open the developer tools, search the bundle for sk-, and copy your key in about a minute.

Once a key is copied, the attacker's usage is billed to you. Suppose a script sends 10 requests a second, each with 20,000 tokens, to the larger model. At $5 per million input tokens, each request costs $0.10, so the script costs $1 a second, or $3,600 an hour. Provider spending limits help, but they also turn off your real feature when they trigger.

What the BFF owns

Hiding the key is only the first reason. Once there is a server in the middle, it becomes the natural place for every decision that must not be controlled by the person holding the browser.

  • The key. Read from a server environment variable, never sent to the client.
  • The prompt. If the browser sent the prompt, any user could rewrite your instructions.
  • The data. The BFF loads reviews by product id from the database. A browser that sent review text could send anything.
  • Limits. Maximum output tokens, maximum words, requests per session.
  • Validation. Output is checked against a schema before it leaves the server.
  • Caching. One summary per product per day is shared by every shopper.
  • Model choice. Switching models is a server deploy, not a new frontend release.
  • Logging. Which prompt version produced which output, for debugging and feedback.

What the browser still owns

A BFF does not make the browser simple. The browser is where time is felt, so it owns the experience of waiting and failing.

  • The state of each request: idle, loading, streaming, done, error, cancelled.
  • Cancelling requests the user no longer wants, with AbortController.
  • Ignoring responses that arrive for a product the user has left.
  • Rendering output safely, without HTML injection.
  • Not asking at all: debouncing input, reusing results already in memory, skipping requests for reviews that are off screen.

A good rule: the server protects money, secrets and truth; the browser protects the user's time and attention.

Design endpoints around features, not around the model

The easiest BFF to write is a proxy: POST /api/llm takes a prompt and returns the model's answer. It hides the key, so it feels safe. It is not. Anyone can call that endpoint with any prompt, and you have built a free, anonymous model API paid for by your company.

TypeScript
// Wrong: a generic proxy. The key is hidden, but anyone can use it.app.post("/api/llm", async (req, res) => {  const text = await llm.complete({ tier: "main", system: "", user: req.body.prompt, maxTokens: 4000 });  res.json({ text });});// Right: one endpoint per feature. The browser sends ids, the server decides the rest.app.post("/api/sentiment", sentimentRoute); // { productId, reviewIds }  -> labelsapp.post("/api/summary", summaryRoute);     // { productId }             -> streamed textapp.post("/api/compare", compareRoute);     // { productIds: [a, b] }     -> table rows

The feature endpoints take small, checkable inputs: an id, a list of up to 20 ids, a pair of ids. The server can reject anything else with a 400. Each endpoint has its own prompt, its own token limit and its own output schema. The worst an abuser can do is ask for summaries of your own products, which you cache anyway.

EndpointRequestResponseModel
POST /api/sentimentproduct id and up to 20 review idsJSON labelssmall
POST /api/summaryproduct idstreamed plain text, max 80 wordslarger
POST /api/comparetwo product ids, optional aspectsJSON rows, max 6larger

Does the extra hop cost time?

The BFF adds one network hop. From the shopper's browser to your server and from your server to the provider is a little slower than going straight to the provider, typically 10 to 40 ms extra. Against a 2,500 ms summary, that is under 2 percent. In exchange, the BFF can return a cached summary in 30 ms, which is far faster than any model call. In practice the BFF makes ShopLens faster, not slower.

Check your understanding

0 of 3 answered

1.The team puts the key in the BFF but exposes POST /api/llm that forwards any prompt. What is the remaining problem?

2.Why does the ShopLens summary endpoint take a product id instead of the review text?

3.Which responsibility belongs in the browser rather than the BFF?