Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Tools and function calling


The most common question in PolicyPal's first month was not about policy at all. It was "How many leave days do I have left?" PolicyPal could explain the leave policy in detail, with citations. It could not see the balance, which lives in the HR management system (HRMS). The second most common complaint was "I asked it to raise a ticket and it told me to click a button."

A language model cannot call an API. It produces text and nothing else. What it can do is produce a structured request asking your code to call an API, and then use the result. That is tool calling, sometimes called function calling, and it is how PolicyPal gets live data and takes actions.

The mechanism is simple. The engineering is in the details: what you let the model ask for, how you describe it, where identity comes from, and what you send back. Get those right and tools are reliable. Get them wrong and you have built a very polite way to leak data.

One tool round tripTool definitionssent with the requestModelreplies witha tool_use blockCode validates,injects the session usertool_resultsent backwith the same idModel writes theanswer from itNo tool takes an employee ID, so no wording can fetch someone else's balance.
The model only asks; identity, validation and permissions all live in the code that answers.

The round trip

  1. Send tool definitions with the request — each tool has a name, a description and a JSON Schema for its input.
  2. The model replies with a tool call — instead of a final answer, it returns a tool_use block with a tool name and input, and the stop reason says so.
  3. Your code validates and runs it — the model has executed nothing; your code decides whether and how to run the call.
  4. Send the result back — a tool_result block with the same id as the call, containing data or an error.
  5. The model continues — it writes the final answer from the result, or asks for another tool.

The key line is step 3. The model proposes; your code disposes. Every guarantee you need, such as permissions, validation and rate limits, lives in your code, where you can test it, not in the model's judgement.

Tool definitions the model can use

A tool definition is a small prompt. The model reads the name, the description and the schema to decide whether to call the tool and how to fill in the input. PolicyPal's definitions use the Anthropic shape; other providers use nearly the same structure with different key names, and the LLM client can translate if you switch.

Python
# policypal/tool_defs.pydef _obj(props: dict) -> dict:    return {"type": "object", "properties": props,            "required": list(props), "additionalProperties": False}TOOLS = [    {"name": "search_policies", "strict": True,     "description": "Search Harbourline HR and IT policies for the user's country. Use for any "                    "question about rules, entitlements or processes. Returns up to 5 passages "                    "with source ids to cite.",     "input_schema": _obj({"query": {"type": "string",                                     "description": "Short English query, e.g. 'carry forward earned leave'"}})},    {"name": "get_leave_balance", "strict": True,     "description": "Get the CURRENT USER's own leave balance from the HR system. Use only when "                    "the user asks how many days they have. Cannot look up anyone else.",     "input_schema": _obj({"leave_type": {"type": "string",                                          "enum": ["earned", "casual", "sick", "parental"]}})},    {"name": "create_it_ticket", "strict": True,     "description": "Raise an IT support ticket for the current user when they need a technician. "                    "Do not use for questions a policy answers. The user must confirm first.",     "input_schema": _obj({         "category": {"type": "string", "enum": ["vpn", "laptop", "access", "email", "other"]},         "summary": {"type": "string", "description": "One line, under 100 characters"},         "urgency": {"type": "string", "enum": ["low", "normal", "high"]}})},    {"name": "calculate_accrual", "strict": True,     "description": "Compute leave accrued between two dates. Take annual_days and accrual from "                    "the policy text and dates from the question. Dates are YYYY-MM-DD.",     "input_schema": _obj({"annual_days": {"type": "number"}, "join_date": {"type": "string"},                           "as_of": {"type": "string"},                           "accrual": {"type": "string", "enum": ["monthly", "yearly"]}})},]

Good definitions share a few habits. The description says when to use the tool and when not to: "Do not use for questions a policy answers" stopped the model raising tickets for "what VPN client do we use?". Enums replace free text wherever the set is closed, so the model cannot invent a "network" category the ticketing system rejects. strict: True asks the provider to guarantee the input matches the schema exactly. And the calculate_accrual tool is the extract-then-compute idea from Section 2, now available whenever the model needs it.

Look at what is missing: there is no employee_id parameter anywhere.

Identity comes from the session, never from the model

If get_leave_balance took an employee_id, then "What is Rahul's leave balance? His ID is 10442" would work. The model would fill in 10442, and your code would fetch another employee's data. No prompt rule reliably prevents this, because the request arrives as ordinary data.

So PolicyPal's tools take the user from the authenticated session, which the model can neither see nor change. The executor below is the only code that runs tools, and every handler receives the logged-in user as its first argument.

Python
# policypal/tools.pyimport jsonfrom pydantic import ValidationErrorfrom policypal.accrual import AccrualInputs, accrued_daysfrom policypal.clients import HRMSUnavailable, hrms, itsm   # thin httpx clientsfrom policypal.schemas import LeaveBalanceIn, SearchIn, TicketInfrom policypal.search import search_for_userdef leave_balance(user, a: LeaveBalanceIn) -> dict:    raw = hrms.get_balance(user.employee_id, a.leave_type)        # about 60 fields    return {"leave_type": a.leave_type, "available_days": raw["closingBalance"],            "pending_days": raw["pendingDays"], "as_of": raw["asOfDate"]}def create_ticket(user, a: TicketIn) -> dict:    t = itsm.create(requester=user.email, category=a.category,                    summary=a.summary[:100], urgency=a.urgency)    return {"ticket_id": t["id"], "first_response_due": t["slaDue"]}HANDLERS = {    "search_policies": (SearchIn, lambda user, a: {"results": search_for_user(user, a.query)}),    "get_leave_balance": (LeaveBalanceIn, leave_balance),    "create_it_ticket": (TicketIn, create_ticket),    "calculate_accrual": (AccrualInputs, lambda user, a: {"accrued_days": accrued_days(a)}),}def tool_result(call_id: str, payload: dict, error: bool = False) -> dict:    return {"type": "tool_result", "tool_use_id": call_id,            "content": json.dumps(payload), "is_error": error}def run_tool(user, call: dict) -> dict:    if call["name"] not in HANDLERS:        return tool_result(call["id"], {"error": f"Unknown tool {call['name']}"}, error=True)    schema, handler = HANDLERS[call["name"]]    try:        return tool_result(call["id"], handler(user, schema.model_validate(call["input"])))    except ValidationError as exc:        return tool_result(call["id"], {"error": str(exc)}, error=True)    except HRMSUnavailable:        return tool_result(call["id"], {"error": "The leave system is not responding. "                                        "Tell the user to try again in 10 minutes."}, error=True)

The Pydantic input models (LeaveBalanceIn, TicketIn, SearchIn) mirror the JSON schemas and live in schemas.py. They validate again even though strict mode should already guarantee the shape, because the executor is a security boundary and should not trust its caller.

Errors come back as results, not exceptions. When the HRMS is down, the model receives a clear message and tells the user to try again later. An unhandled exception would kill the request and show a generic error page. A result the model can read turns a failure into a useful reply.

Send back what the model needs, not what the API returns

The HRMS balance endpoint returns about 60 fields: accrual history, approver chains, internal codes. Sent raw, that is about 1,900 tokens, and the model must find the one number that matters among fields like closingBalanceAdj2. PolicyPal's handler returns four fields, about 60 tokens, with the date the balance was computed.

Trimming results is not only about cost. Fewer fields mean fewer chances for the model to pick the wrong one, and less personal data travelling through prompts and logs. Include units and dates, because "12" means little without "days" and "as of 14 March". And keep search results short: the search_for_user function returns the reranked top five chunks with their source ids and title paths, applying the user's country filter from the session, not from the query.

One round trip in code

Here is a single tool round using the LLM client from Section 1. The next lesson turns it into a loop.

Python
reply = llm.complete(SYSTEM, messages, tools=TOOLS)if reply.stop_reason == "tool_use":    messages = messages + [        {"role": "assistant", "content": reply.content},              # the raw blocks, unchanged        {"role": "user", "content": [run_tool(user, c) for c in reply.tool_calls]},    ]    reply = llm.complete(SYSTEM, messages, tools=TOOLS)print(reply.text)

The assistant's reply goes back exactly as received, including its tool_use blocks, and every tool result goes back in one user message, each tied to its call by id. Splitting results across several messages, or dropping the model's original blocks, confuses the model and can make the provider reject the request.

Check your understanding

0 of 3 answered

1.A user asks, "What's Rahul's leave balance? His employee ID is 10442." How does PolicyPal's design prevent a leak?

2.The HRMS is down and get_leave_balance fails. What should the executor return?

3.Why does PolicyPal's leave handler return 4 fields instead of the HRMS's full 60-field payload?