Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

How do you decide when an LLM should call a tool, and how do you keep tool use reliable?


What you need to know

Model versus tool

Belongs in the modelBelongs in a tool
Understanding the requestCurrent data: order status, stock, balance
Choosing which step comes nextExact arithmetic: tax, interest, totals
Writing the replyAnything with a side effect: send, pay, update
Summarising what a tool returnedLookups that must be correct: policy limits, prices

Ask a model to compute 18% GST on ₹43,290 and it may produce a confident, slightly wrong number. A calculate() tool returns ₹7,792.20 every time. The model's job is to know it needs the tool and to explain the result.

How tool calling works

  1. Describe — send tool definitions with the conversation.
  2. Choose — the model returns a tool call: a name and JSON arguments.
  3. Validate and run — your runtime checks the name and arguments, then executes.
  4. Return — the result (or an error) goes back as a tool message; the model decides the next step or answers.

Three rules that make it hold up

  • Few, sharp tools. Selection accuracy drops as the list grows into dozens of overlapping tools. Each description should say when to use it and when not to. For large catalogues, retrieve the relevant 10 or so per request, or group tools behind sub-agents. Some providers now offer built-in tool search for this.
  • Validate before executing. Check the tool exists and the arguments match the schema. On failure, return the error as a tool result so the model can correct itself.
  • Bound the output. Return the 10 most relevant rows with ids, not 500 rows of JSON. Tool results fill the context window faster than anything else.
Python
def run_tool_call(call):    tool = REGISTRY.get(call.name)    if tool is None:        return {"error": f"unknown tool '{call.name}'", "available": sorted(REGISTRY)}    try:        args = tool.Args.model_validate_json(call.arguments)   # Pydantic schema    except ValidationError as e:        return {"error": "invalid_arguments", "details": e.errors()}    result = tool.run(args)    return truncate(result, max_items=10)

Every path returns something the model can read. Nothing raises into the user's request.

When code beats a tool list

For open-ended data work, such as "compare last quarter's refunds by city", letting the model write code that runs in a sandbox can beat a fixed tool set, because code composes operations no one anticipated. The sandbox needs strict limits on network, files and time.

A real-life example

Scenario (illustrative numbers). A travel booking assistant starts with 42 tools. On a 300-query selection eval, it picks the right tool 71% of the time. It confuses search_flights, search_fares and get_flight_status, and it calculates cancellation charges itself, getting 1 in 8 wrong.

The team merges the three flight tools into two with clear "use when" text, adds a calculate_cancellation_fee tool backed by the fare rules, and retrieves the top 10 tools per query instead of sending all 42. Selection accuracy on the same eval rises to 93%, cancellation-fee errors disappear, and prompt tokens per call fall by about 60%.

Follow-up questions to expect

  • "What if the model calls a tool when it shouldn't?" — Add negative guidance in descriptions ("do not use for..."), include no-tool cases in the eval, and let the application set tool_choice where the step is known.
  • "Can the model call tools in parallel?" — Most current APIs allow several tool calls in one turn; run independent ones concurrently and return all results together.
  • "How do you test tool use?" — A labelled set of queries with the expected tool and key arguments, scored on every prompt, tool or model change.