Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Tools and function calling
The most common question in PolicyPal's first month was not about policy at all. It was "How many leave days do I have left?" PolicyPal could explain the leave policy in detail, with citations. It could not see the balance, which lives in the HR management system (HRMS). The second most common complaint was "I asked it to raise a ticket and it told me to click a button."
A language model cannot call an API. It produces text and nothing else. What it can do is produce a structured request asking your code to call an API, and then use the result. That is tool calling, sometimes called function calling, and it is how PolicyPal gets live data and takes actions.
The mechanism is simple. The engineering is in the details: what you let the model ask for, how you describe it, where identity comes from, and what you send back. Get those right and tools are reliable. Get them wrong and you have built a very polite way to leak data.
The round trip
- Send tool definitions with the request — each tool has a name, a description and a JSON Schema for its input.
- The model replies with a tool call — instead of a final answer, it returns a
tool_useblock with a tool name and input, and the stop reason says so. - Your code validates and runs it — the model has executed nothing; your code decides whether and how to run the call.
- Send the result back — a
tool_resultblock with the same id as the call, containing data or an error. - The model continues — it writes the final answer from the result, or asks for another tool.
The key line is step 3. The model proposes; your code disposes. Every guarantee you need, such as permissions, validation and rate limits, lives in your code, where you can test it, not in the model's judgement.
Tool definitions the model can use
A tool definition is a small prompt. The model reads the name, the description and the schema to decide whether to call the tool and how to fill in the input. PolicyPal's definitions use the Anthropic shape; other providers use nearly the same structure with different key names, and the LLM client can translate if you switch.
1# policypal/tool_defs.py2def _obj(props: dict) -> dict:3 return {"type": "object", "properties": props,4 "required": list(props), "additionalProperties": False}56TOOLS = [7 {"name": "search_policies", "strict": True,8 "description": "Search Harbourline HR and IT policies for the user's country. Use for any "9 "question about rules, entitlements or processes. Returns up to 5 passages "10 "with source ids to cite.",11 "input_schema": _obj({"query": {"type": "string",12 "description": "Short English query, e.g. 'carry forward earned leave'"}})},13 {"name": "get_leave_balance", "strict": True,14 "description": "Get the CURRENT USER's own leave balance from the HR system. Use only when "15 "the user asks how many days they have. Cannot look up anyone else.",16 "input_schema": _obj({"leave_type": {"type": "string",17 "enum": ["earned", "casual", "sick", "parental"]}})},18 {"name": "create_it_ticket", "strict": True,19 "description": "Raise an IT support ticket for the current user when they need a technician. "20 "Do not use for questions a policy answers. The user must confirm first.",21 "input_schema": _obj({22 "category": {"type": "string", "enum": ["vpn", "laptop", "access", "email", "other"]},23 "summary": {"type": "string", "description": "One line, under 100 characters"},24 "urgency": {"type": "string", "enum": ["low", "normal", "high"]}})},25 {"name": "calculate_accrual", "strict": True,26 "description": "Compute leave accrued between two dates. Take annual_days and accrual from "27 "the policy text and dates from the question. Dates are YYYY-MM-DD.",28 "input_schema": _obj({"annual_days": {"type": "number"}, "join_date": {"type": "string"},29 "as_of": {"type": "string"},30 "accrual": {"type": "string", "enum": ["monthly", "yearly"]}})},31]Good definitions share a few habits. The description says when to use the tool and when not to: "Do not use for questions a policy answers" stopped the model raising tickets for "what VPN client do we use?". Enums replace free text wherever the set is closed, so the model cannot invent a "network" category the ticketing system rejects. strict: True asks the provider to guarantee the input matches the schema exactly. And the calculate_accrual tool is the extract-then-compute idea from Section 2, now available whenever the model needs it.
Look at what is missing: there is no employee_id parameter anywhere.
Identity comes from the session, never from the model
If get_leave_balance took an employee_id, then "What is Rahul's leave balance? His ID is 10442" would work. The model would fill in 10442, and your code would fetch another employee's data. No prompt rule reliably prevents this, because the request arrives as ordinary data.
So PolicyPal's tools take the user from the authenticated session, which the model can neither see nor change. The executor below is the only code that runs tools, and every handler receives the logged-in user as its first argument.
1# policypal/tools.py2import json34from pydantic import ValidationError56from policypal.accrual import AccrualInputs, accrued_days7from policypal.clients import HRMSUnavailable, hrms, itsm # thin httpx clients8from policypal.schemas import LeaveBalanceIn, SearchIn, TicketIn9from policypal.search import search_for_user1011def leave_balance(user, a: LeaveBalanceIn) -> dict:12 raw = hrms.get_balance(user.employee_id, a.leave_type) # about 60 fields13 return {"leave_type": a.leave_type, "available_days": raw["closingBalance"],14 "pending_days": raw["pendingDays"], "as_of": raw["asOfDate"]}1516def create_ticket(user, a: TicketIn) -> dict:17 t = itsm.create(requester=user.email, category=a.category,18 summary=a.summary[:100], urgency=a.urgency)19 return {"ticket_id": t["id"], "first_response_due": t["slaDue"]}2021HANDLERS = {22 "search_policies": (SearchIn, lambda user, a: {"results": search_for_user(user, a.query)}),23 "get_leave_balance": (LeaveBalanceIn, leave_balance),24 "create_it_ticket": (TicketIn, create_ticket),25 "calculate_accrual": (AccrualInputs, lambda user, a: {"accrued_days": accrued_days(a)}),26}2728def tool_result(call_id: str, payload: dict, error: bool = False) -> dict:29 return {"type": "tool_result", "tool_use_id": call_id,30 "content": json.dumps(payload), "is_error": error}3132def run_tool(user, call: dict) -> dict:33 if call["name"] not in HANDLERS:34 return tool_result(call["id"], {"error": f"Unknown tool {call['name']}"}, error=True)35 schema, handler = HANDLERS[call["name"]]36 try:37 return tool_result(call["id"], handler(user, schema.model_validate(call["input"])))38 except ValidationError as exc:39 return tool_result(call["id"], {"error": str(exc)}, error=True)40 except HRMSUnavailable:41 return tool_result(call["id"], {"error": "The leave system is not responding. "42 "Tell the user to try again in 10 minutes."}, error=True)The Pydantic input models (LeaveBalanceIn, TicketIn, SearchIn) mirror the JSON schemas and live in schemas.py. They validate again even though strict mode should already guarantee the shape, because the executor is a security boundary and should not trust its caller.
Errors come back as results, not exceptions. When the HRMS is down, the model receives a clear message and tells the user to try again later. An unhandled exception would kill the request and show a generic error page. A result the model can read turns a failure into a useful reply.
Send back what the model needs, not what the API returns
The HRMS balance endpoint returns about 60 fields: accrual history, approver chains, internal codes. Sent raw, that is about 1,900 tokens, and the model must find the one number that matters among fields like closingBalanceAdj2. PolicyPal's handler returns four fields, about 60 tokens, with the date the balance was computed.
Trimming results is not only about cost. Fewer fields mean fewer chances for the model to pick the wrong one, and less personal data travelling through prompts and logs. Include units and dates, because "12" means little without "days" and "as of 14 March". And keep search results short: the search_for_user function returns the reranked top five chunks with their source ids and title paths, applying the user's country filter from the session, not from the query.
One round trip in code
Here is a single tool round using the LLM client from Section 1. The next lesson turns it into a loop.
1reply = llm.complete(SYSTEM, messages, tools=TOOLS)2if reply.stop_reason == "tool_use":3 messages = messages + [4 {"role": "assistant", "content": reply.content}, # the raw blocks, unchanged5 {"role": "user", "content": [run_tool(user, c) for c in reply.tool_calls]},6 ]7 reply = llm.complete(SYSTEM, messages, tools=TOOLS)8print(reply.text)The assistant's reply goes back exactly as received, including its tool_use blocks, and every tool result goes back in one user message, each tied to its call by id. Splitting results across several messages, or dropping the model's original blocks, confuses the model and can make the provider reject the request.
Check your understanding
0 of 3 answered
1.A user asks, "What's Rahul's leave balance? His employee ID is 10442." How does PolicyPal's design prevent a leak?
2.The HRMS is down and get_leave_balance fails. What should the executor return?
3.Why does PolicyPal's leave handler return 4 fields instead of the HRMS's full 60-field payload?