Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your LLM invents a function name that doesn't exist — get_user_full_history() — and your runtime crashes 0.5% of the time. How do you make tool calling crash-proof in production?


What you need to know

The crash happens because the runtime trusts the model. Code like getattr(tools, name)(**args) raises the moment the model invents get_user_full_history. The model is not wrong in a surprising way; it guessed a plausible capability. Your runtime must expect that.

The dispatch path

  1. Look up, never invoke dynamically — a dictionary registry; an unknown name returns an error object.
  2. Validate arguments — Pydantic model per tool; missing fields, wrong types and made-up enum values become readable errors.
  3. Execute in a guard — timeout, retry policy for transient failures, and a catch-all that converts any exception into an error result while logging the traceback.
  4. Bound retries — two failed attempts on the same tool end the run with a graceful message.
Python
def dispatch(name: str, raw_args: str) -> dict:    tool = REGISTRY.get(name)    if tool is None:        return {"error": "unknown_tool", "message": f"'{name}' does not exist.",                "available": sorted(REGISTRY)[:20]}    try:        args = tool.Args.model_validate_json(raw_args)    except ValidationError as e:        return {"error": "invalid_arguments", "details": e.errors(include_url=False)}    try:        return run_with_timeout(tool.run, args, seconds=10)    except Exception:        log.exception("tool_failed", tool=name)        return {"error": "tool_failed", "message": "The tool failed. Try another approach."}

Every branch returns data the model can read. The traceback goes to your logs, not into the model's context, where it would waste tokens and could leak internals.

Why returning errors beats raising

ApproachWhat the user seesWhat the model can do
Raise an exceptionA 500 errorNothing; the run is over
Return a structured errorA slightly slower answerPick a real tool, fix the arguments, or explain

Then reduce the rate

Hallucinated names are data. Group them from the logs: they usually cluster around a real capability you don't have, or a real tool whose name doesn't match how the model thinks about the task.

  • If the model keeps asking for get_user_full_history, maybe users need it: add a real list_user_orders.
  • If it asks for search_customer when you have lookup_account, rename or describe the tool in the words the model uses.

Current APIs also offer strict schema modes for tool arguments, which prevent many argument-shape errors. They don't stop an invented tool name in every setup, so the registry check stays.

Metrics

Unknown-tool rate, validation-failure rate, and recovery rate: the share of runs that succeed after an error result. All three belong on a dashboard.

A real-life example

Scenario (illustrative numbers). A telecom's customer-care agent handles 400,000 conversations a month. About 2,000 end in a 500 error because the model calls tools that don't exist, most often get_user_full_history and check_network_outage.

The team switches to registry dispatch with structured errors. Crashes drop to zero, and the recovery rate after an unknown-tool error is 89%: the model picks get_recent_tickets instead and continues. The logs show check_network_outage was requested 1,300 times in a month, so the team builds that tool, backed by the network-status API. Unknown-tool errors fall by 70%.

Follow-up questions to expect

  • "Should you show the full list of tools in the error?" — A short list of relevant tools helps; the full catalogue wastes tokens. Twenty names or the domain's tools are enough.
  • "What about a tool that hangs?" — A per-tool timeout, then an error result; the agent's overall deadline is the backstop.
  • "Is this different in a framework like LangGraph?" — The same principle; frameworks offer error-handling options for tool nodes, but you should still configure what the model sees on failure.