Agents & Tools Interview Prep

Course Content

Agents & Tools Interview Prep

6 sections · 40 lessons

What are best practices when defining tools (inputs, outputs, descriptions)?


The travel tool that failed one call in fiveOne travel_api tool• op: free-text string• params: any object• Description: Travel operations• 60 fields per hotel returnedFour typed tools• search_flights, search_hotels ...• Formats, enums, strict mode on• Says when to use and when not• 7 fields, 10 results, a page cursor
The model only knows what the definition tells it, so every vague field becomes a guess at call time.

What you need to know

Names

  • Verb first, specific, snake_case: search_flights, get_booking, cancel_booking.
  • One tool, one job. Avoid manage_booking(action=...) with ten modes — the model must then get both the tool and the mode right.
  • Group related tools with a prefix (github_list_issues, github_add_label), especially when tools come from several sources.

Descriptions

Say what, when, when not, and what comes back:

Text
Search available hotels for given dates. Use when the user wants hotel optionsor prices. Do not use to check an existing reservation - use get_booking.Returns at most 10 hotels sorted by price, each with hotel_id, name, area,price_per_night_inr and rating.

Current Claude models are fairly conservative about calling tools, and trigger conditions ("use when…") in the description measurably improve how often they call the right one.

Input schemas

ProblemWeak schemaBetter schema
Unclear format"date": {"type": "string"}"format": "date", description "YYYY-MM-DD, e.g. 2026-12-20"
Unclear units"budget": {"type": "number"}"budget_inr" with "total for the stay, in rupees"
Free text for a fixed set"cabin": {"type": "string"}"enum": ["economy", "premium_economy", "business"]
Invented IDs"hotel_id" with no hint"hotel_id from search_hotels results, e.g. HTL-4411"
Extra junk fieldsnothing"additionalProperties": false

Strict mode

Constrained decoding makes the model's arguments always match the schema.

  • Anthropic: add "strict": true at the top level of the tool definition (next to name), with additionalProperties: false and a required list.
  • OpenAI: "strict": true on the function tool. Every property must be listed in required; an optional field is written as nullable, e.g. "type": ["string", "null"].

Strict mode fixes shape errors (missing fields, wrong types, bad enum values). It does not fix meaning errors (a real but wrong date), so you still validate in the handler.

Outputs

  • Return only the fields the model needs, not a 40-field database row.
  • Cap and paginate: "10 results; call again with page=2 for more".
  • Use stable IDs the model can pass to the next tool.
  • Prefer structured JSON or clean text over raw HTML.

Errors

Make errors instructive: "No hotel with id HTL-44. Use a hotel_id from search_hotels results." beats "404". The model reads the error and usually fixes the call on the next turn.

How many tools

Accuracy falls as the menu grows and overlaps. If you have dozens of tools, load them by task, merge near-duplicates, or use tool search: mark rarely used tools with defer_loading: true so the model searches for them instead of reading every schema up front (both Anthropic and OpenAI support this in 2026).

A real-life example

A travel startup's first version had one tool:

JSON
{"name": "travel_api", "description": "Travel operations", "input_schema": {"type": "object", "properties": {   "op": {"type": "string"}, "params": {"type": "object"}}}}

The model guessed operation names ("find_flight", "flightSearch"), sent dates as "20/12", and invented hotel IDs. About 1 in 5 calls failed.

They split it into search_flights, search_hotels, get_booking and hold_booking, each with typed parameters, formats, enums and "use when" descriptions, turned on strict mode, and trimmed hotel results from 60 fields to 7. Failed calls dropped to under 1 in 50, and the average tool result shrank from about 9,000 tokens to 900 — which also cut cost per conversation by more than half.

Follow-up questions to expect

  • "Should you give examples inside tool definitions?" — Yes, for complex inputs. A format example in the parameter description is cheap; Anthropic also supports example inputs on the tool definition.
  • "Should a tool return errors as exceptions or results?" — As results with an error flag, so the model can adapt. Exceptions should only escape for bugs in your harness.
  • "How do you know your descriptions are good?" — Build a labelled set of requests with the correct tool and arguments, and measure tool-call accuracy after each change.