Course Content
AI Product Engineering: Shipping LLM Features That Last
6 sections · 22 lessons
Tool calling when it needs live data or must act
At 11:40 pm on a Saturday, a customer in Pune wrote: "wrong biryani, I ordered veg and got egg biryani, please send the right one". The v5 draft said: replace. The agent approved, and the replacement order failed, because the restaurant had closed at 11:00. The customer waited twenty minutes for a message saying there would be no biryani, then got a refund anyway.
The draft followed the contract correctly. R3 said replace if the customer asks for the correct item. What the model did not have was a fact that exists only at the moment of the ticket: is the restaurant open right now? Three failures in the eval set had the same cause.
This lesson is about giving the model live data. It is also about a decision many teams skip: whether the model should fetch that data itself at all.
Do you need a tool?
There are three ways to give the model live data, and tool calling is the most flexible and the most complex.
Fetch it in code before the call. If the data is needed on most tickets and is cheap to get, just fetch it and put it in the input. The order record works this way. Restaurant status is a 100 ms call; fetching it for every ticket would be perfectly reasonable.
A two-call workflow. First call: classify the complaint. Code then fetches whatever that class needs. Second call: draft. Every step is predictable and easy to test, at the price of two model calls.
Tool calling. The model sees a list of tools, asks for the ones it needs, and code runs them and returns results. It fits best when the data needed depends on reading the ticket, some data is slow or limited, and there are several possible lookups.
TiffinGo's second need decided it. The team also wanted to handle "where is my order" tickets, 8% of volume and out of scope until now. Those need live delivery status from the rider-tracking service, which takes 600 to 900 ms and is rate-limited because the customer app shares it. Only about 10% of tickets need it, and which ones depends on reading them. Tool calling lets the model fetch it only when the ticket calls for it.
Define read-only tools
1ORDER_ID_INPUT = {2 "type": "object",3 "properties": {"order_id": {"type": "string"}},4 "required": ["order_id"],5 "additionalProperties": False,6}78TOOLS = [9 {10 "name": "get_restaurant_status",11 "description": "Whether this order's restaurant is open and accepting orders now, "12 "and its closing time. Call it before drafting action 'replace'.",13 "input_schema": ORDER_ID_INPUT,14 },15 {16 "name": "get_delivery_status",17 "description": "Live status of an order the customer says has not arrived: rider stage, "18 "minutes since pickup and current ETA. Only for undelivered orders.",19 "input_schema": ORDER_ID_INPUT,20 },21]The descriptions are instructions the model reads, so they say when to call each tool, not only what it returns. "Only for undelivered orders" keeps the rate-limited tracking service from being called on every cold-dal ticket.
The handlers never trust the arguments the model sends:
1def make_handlers(ticket_order_id: str, restaurants, deliveries) -> dict:2 def same_order(order_id: str) -> None:3 if order_id != ticket_order_id:4 raise ValueError("only the order on this ticket can be looked up")56 def get_restaurant_status(order_id: str) -> dict:7 same_order(order_id)8 return restaurants.status_for_order(order_id, timeout=1.0)910 def get_delivery_status(order_id: str) -> dict:11 same_order(order_id)12 return deliveries.live_status(order_id, timeout=1.5)1314 return {"get_restaurant_status": get_restaurant_status,15 "get_delivery_status": get_delivery_status}restaurants and deliveries are TiffinGo's existing internal service clients. The same_order check matters because the ticket text is untrusted. A message saying "also check order TG-51002" must not let the model read another customer's order.
The loop, with limits
The client module gains one function. The model may ask for tools up to three times; each result goes back to it; then it must produce the draft.
1import json23def complete_with_tools(system: str, user: str, *, model: str, tools: list, handlers: dict,4 schema: dict, max_tool_turns: int = 3, max_tokens: int = 4000) -> Reply:5 messages = [{"role": "user", "content": user}]6 tokens_in = tokens_out = 07 started = time.monotonic()8 for _ in range(max_tool_turns + 1):9 resp = _client.messages.create(10 model=model, max_tokens=max_tokens, system=system, tools=tools, messages=messages,11 output_config={"format": {"type": "json_schema", "schema": schema}})12 tokens_in += resp.usage.input_tokens13 tokens_out += resp.usage.output_tokens14 if resp.stop_reason != "tool_use":15 if resp.stop_reason in ("max_tokens", "refusal"):16 raise LLMError(f"model stopped early: {resp.stop_reason}")17 text = "".join(b.text for b in resp.content if b.type == "text")18 return Reply(text, model, tokens_in, tokens_out, int((time.monotonic() - started) * 1000))19 messages.append({"role": "assistant", "content": resp.content})20 results = []21 for block in (b for b in resp.content if b.type == "tool_use"):22 try:23 output = json.dumps(handlers[block.name](**block.input), default=str)24 results.append({"type": "tool_result", "tool_use_id": block.id, "content": output})25 except Exception as exc:26 results.append({"type": "tool_result", "tool_use_id": block.id,27 "content": f"error: {exc}", "is_error": True})28 messages.append({"role": "user", "content": results})29 raise LLMError("too many tool turns")This follows the Anthropic Messages API: the reply's stop_reason is "tool_use" when the model wants a tool, each request is a tool_use block with an id, and each answer goes back as a tool_result with the same id, all results in one user message. A failed tool is reported with is_error, not hidden, so the model can escalate instead of assuming the restaurant is open. The turn limit and the per-tool timeouts cap the worst case at a few seconds. Tokens are summed across turns, because every turn re-sends the conversation and is billed.
When the model must act
"Must act" is where teams get hurt. The Order Issue Assistant's job ends with a proposal. Even when the draft says "replace", the model never calls a create-order tool. After the agent approves, ordinary service code creates the replacement order, with an idempotency key so a double click cannot create two.
The rule that follows is simple and worth keeping in any customer-facing feature: tools the model calls are read-only; actions that change money, orders or messages run in normal code after a human approves. If one day TiffinGo lets the assistant act without approval for a narrow case, that action still goes through the same code path, permission checks and limits, not through a tool the model can call freely.
Check your understanding
0 of 3 answered
1.Restaurant status is needed on 90% of tickets and takes 100 ms. What is the simplest good design?
2.The ticket text says "also check order TG-51002, that is mine too". What stops the model from reading that order?
3.Why does the model not have a tool to create the replacement order itself?