Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 5: Custom Tool-Oriented Agent Design
Scenario: the team wants a tool-using agent in plain Python, with no LangChain or CrewAI, because they want to own the loop. How do you design it for production?
What you need to know
The loop
1TOOLS = {"get_order": GetOrderTool(), "get_policy": GetPolicyTool()} # name -> tool23def run_agent(messages, max_steps=8, deadline_s=45):4 start, seen = time.monotonic(), set()5 for _ in range(max_steps):6 if time.monotonic() - start > deadline_s:7 return "Sorry, this is taking too long. A teammate will follow up."8 msg = client.chat.completions.create(9 model=MODEL, messages=messages,10 tools=[t.schema for t in TOOLS.values()]).choices[0].message11 messages.append(msg)12 if not msg.tool_calls:13 return msg.content14 for call in msg.tool_calls:15 result = dispatch(call, seen) # never raises16 messages.append({"role": "tool", "tool_call_id": call.id,17 "content": json.dumps(result)[:4000]})18 return "I couldn't finish this. Escalating to a human."This uses the OpenAI Chat Completions tool-calling format; other providers use the same pattern with different field names.
What dispatch does
1def dispatch(call, seen):2 tool = TOOLS.get(call.function.name)3 if tool is None:4 return {"error": f"unknown tool '{call.function.name}'", "available": list(TOOLS)}5 sig = (call.function.name, call.function.arguments)6 if sig in seen:7 return {"error": "repeat_call", "hint": "You already have this result. Use it or answer."}8 seen.add(sig)9 try:10 args = tool.Args.model_validate_json(call.function.arguments)11 return tool.run(args)12 except ValidationError as e:13 return {"error": "invalid_arguments", "details": e.errors()}14 except Exception:15 log.exception("tool failed", tool=call.function.name)16 return {"error": "tool_failed"}Every path returns a dict the model can read. The model recovers from a readable error; a raised exception just gives the user a 500.
The design decisions to defend
| Decision | Why |
|---|---|
Registry, not eval or getattr | A hallucinated name cannot run arbitrary code |
| Pydantic validation before running | A made-up field or wrong type never reaches the database |
| Truncated tool results | Large results blow the context window and the bill |
| Step cap, deadline, repeat detector | Guarantees the run ends |
| Writes return a pending action | A human approves anything with side effects |
| A trace id on every step | You can replay any run |
- Read-only tools — ship lookups first and measure success rate.
- Tracing — log every step: model call, tool, arguments, result size, latency.
- Pending writes — side-effecting tools create an approval record instead of acting.
- Metrics — task success rate, average steps, cost per task, abort reasons.
A real-life example
Scenario (illustrative numbers). A D2C skincare brand builds an order-support agent in about 200 lines of Python with four read-only tools. In the first week of internal testing, 3% of runs crash: the model calls get_order_history, which doesn't exist, and a getattr lookup raises.
The team switches to a registry with readable errors, Pydantic validation and a repeat detector. Crashes go to zero. Traces show the model self-corrects to get_order within one step in 92% of unknown-tool cases. Average steps per task is 2.6, and 0.8% of runs hit the 8-step cap, which the team reviews weekly to find missing tools. After a month they add create_return_request as a pending-approval tool.
Follow-up questions to expect
- "Why not use a framework?" — Frameworks add checkpointing, streaming and interrupts for free; owning the loop gives full control and fewer dependencies. Both are valid; the safety controls are the same either way.
- "How do you run independent tool calls faster?" — Execute them concurrently with
asyncio.gatherwhen the model returns several in one turn. - "How do you add memory across sessions?" — Persist the message list, or a summary plus key facts, keyed by conversation id.