Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is tool use and function calling in LLMs, and why is it essential for agents?


One turn of native tool callingModel returnsstructured tool callsCode validatesargs and permissionsThree read-onlycalls run in parallelCapped resultsreturn as tool messagesModel fixes a badargument, goes onThe model proposes; your code decides whether anything runs.
An error string the model can read is a feature — it turns a failed call into a corrected one on the next step.

What you need to know

How a tool call works

  1. You send the model the conversation plus tool definitions.
  2. The model returns either text (final answer) or one or more tool calls: tool name plus JSON arguments.
  3. Your code validates the arguments, checks permissions, runs the function.
  4. You append the result as a tool message and call the model again.

The model never runs anything itself. Your code is always in the middle, which is where you enforce safety.

JSON
{  "name": "get_claim",  "description": "Fetch one insurance claim by ID. Use before any decision on a claim.",  "input_schema": {    "type": "object",    "properties": {"claim_id": {"type": "string", "pattern": "^CLM-[0-9]{4,8}$"}},    "required": ["claim_id"]  }}

ReAct, then and now

ReAct (reason plus act, 2022) had the model write "Thought: ... Action: ... Observation: ..." as text, and code parsed it. Today, providers train models to emit tool calls as structured output, so parsing errors mostly disappear. The idea of ReAct, think then act then observe, is still the loop every agent runs. Reasoning models can now also think between tool calls, which improves multi-step choices.

A framework-neutral loop

Python
import jsondef run_agent(llm, tools, messages, max_steps=12):    for step in range(max_steps):        reply = llm(messages, tools=list(tools))          # model decides        messages.append({"role": "assistant", "content": reply})        calls = [c for c in reply if c["type"] == "tool_call"]        if not calls:                                      # no tool call = final answer            return reply[0]["text"]        for call in calls:                                 # may be several, in parallel            fn = tools.get(call["name"])            try:                result = fn(**call["args"]) if fn else f"unknown tool {call['name']}"            except Exception as e:                result = f"error: {e}"                     # errors go back to the model            messages.append({"role": "tool", "id": call["id"],                             "content": json.dumps(result)[:4000]})    return "Stopped: step budget reached; escalating to a human."

Three details matter. The step budget is in code. Tool errors are returned to the model as text, so it can correct itself. And each result is capped at 4,000 characters so one tool cannot flood the context. In production you would also validate call["args"] against the schema before running.

What makes tools work well

  • Few tools with clear, non-overlapping purposes. Selection accuracy drops as tools multiply.
  • Enums instead of free strings where possible.
  • Error messages the model can act on: "claim_id must look like CLM-1234", not a stack trace.
  • Authorisation in the tool, using the end user's identity.
  • Idempotency keys on writes, so a retry does not act twice.

Where MCP fits

MCP (Model Context Protocol) is an open protocol for exposing tools, data and prompt templates from separate servers to any compatible agent host. The agent host lists a server's tools, shows them to the model as ordinary tool definitions, and forwards calls. It standardises the plumbing; it does not change how the model decides.

A real-life example

A DevOps incident-triage agent gets an alert: "checkout-service error rate 12%". In one turn the model issues three tool calls in parallel: get_recent_deploys(service="checkout"), query_metrics(service="checkout", window="30m") and search_logs(service="checkout", level="error"). The three results come back in about 1.5 seconds together instead of about 4 seconds one after another.

The logs tool returns an error: "window must be one of 15m, 1h, 6h". The model corrects to "1h" on the next step. With the deploy and log evidence together, it posts: "Deploy 4f2a at 14:02 changed the payments client timeout; errors began 14:03." A human decides on the rollback.

Follow-up questions to expect

  • "What if the model calls a tool that does not exist or passes bad arguments?" — Validate against the schema before running, return a clear error, and let the model retry. Count these errors per tool; a high rate means the description is unclear.
  • "How many tools can an agent handle?" — It depends on the model, but accuracy falls as tools overlap. Past a few dozen, load only the relevant tools per task, or split into sub-agents.
  • "Is structured output the same as tool calling?" — Closely related. Both constrain the model to a schema. Tool calling adds the loop: your code runs something and returns a result.