Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 5: Custom Tool-Oriented Agent Design


Scenario: the team wants a tool-using agent in plain Python, with no LangChain or CrewAI, because they want to own the loop. How do you design it for production?

What you need to know

The loop

Python
TOOLS = {"get_order": GetOrderTool(), "get_policy": GetPolicyTool()}   # name -> tooldef run_agent(messages, max_steps=8, deadline_s=45):    start, seen = time.monotonic(), set()    for _ in range(max_steps):        if time.monotonic() - start > deadline_s:            return "Sorry, this is taking too long. A teammate will follow up."        msg = client.chat.completions.create(            model=MODEL, messages=messages,            tools=[t.schema for t in TOOLS.values()]).choices[0].message        messages.append(msg)        if not msg.tool_calls:            return msg.content        for call in msg.tool_calls:            result = dispatch(call, seen)                 # never raises            messages.append({"role": "tool", "tool_call_id": call.id,                             "content": json.dumps(result)[:4000]})    return "I couldn't finish this. Escalating to a human."

This uses the OpenAI Chat Completions tool-calling format; other providers use the same pattern with different field names.

What dispatch does

Python
def dispatch(call, seen):    tool = TOOLS.get(call.function.name)    if tool is None:        return {"error": f"unknown tool '{call.function.name}'", "available": list(TOOLS)}    sig = (call.function.name, call.function.arguments)    if sig in seen:        return {"error": "repeat_call", "hint": "You already have this result. Use it or answer."}    seen.add(sig)    try:        args = tool.Args.model_validate_json(call.function.arguments)        return tool.run(args)    except ValidationError as e:        return {"error": "invalid_arguments", "details": e.errors()}    except Exception:        log.exception("tool failed", tool=call.function.name)        return {"error": "tool_failed"}

Every path returns a dict the model can read. The model recovers from a readable error; a raised exception just gives the user a 500.

The design decisions to defend

DecisionWhy
Registry, not eval or getattrA hallucinated name cannot run arbitrary code
Pydantic validation before runningA made-up field or wrong type never reaches the database
Truncated tool resultsLarge results blow the context window and the bill
Step cap, deadline, repeat detectorGuarantees the run ends
Writes return a pending actionA human approves anything with side effects
A trace id on every stepYou can replay any run
  1. Read-only tools — ship lookups first and measure success rate.
  2. Tracing — log every step: model call, tool, arguments, result size, latency.
  3. Pending writes — side-effecting tools create an approval record instead of acting.
  4. Metrics — task success rate, average steps, cost per task, abort reasons.

A real-life example

Scenario (illustrative numbers). A D2C skincare brand builds an order-support agent in about 200 lines of Python with four read-only tools. In the first week of internal testing, 3% of runs crash: the model calls get_order_history, which doesn't exist, and a getattr lookup raises.

The team switches to a registry with readable errors, Pydantic validation and a repeat detector. Crashes go to zero. Traces show the model self-corrects to get_order within one step in 92% of unknown-tool cases. Average steps per task is 2.6, and 0.8% of runs hit the 8-step cap, which the team reviews weekly to find missing tools. After a month they add create_return_request as a pending-approval tool.

Follow-up questions to expect

  • "Why not use a framework?" — Frameworks add checkpointing, streaming and interrupts for free; owning the loop gives full control and fewer dependencies. Both are valid; the safety controls are the same either way.
  • "How do you run independent tool calls faster?" — Execute them concurrently with asyncio.gather when the model returns several in one turn.
  • "How do you add memory across sessions?" — Persist the message list, or a summary plus key facts, keyed by conversation id.